Contents
Map

02 · Prog Langs

Python for AI Engineering

View as:

Python for AI Engineering

Most AI engineering is ordinary software engineering in Python: managing environments and dependencies, writing typed and tested code, and handling many slow network calls to models at once. This note covers the parts that matter most for LLM work - project setup with uv and pyproject.toml, type hints and Pydantic models, asyncio for concurrent model calls with rate limits, timeouts and retries, testing code that calls non-deterministic models, and profiling.

Learning objectives 50 min
By the end of this page you will be able to:
  • Set up a reproducible Python project with pyproject.toml, a lockfile and dependency groups
  • Use type hints and Pydantic models to validate data at the boundaries of an LLM application
  • Run many model calls concurrently with asyncio while respecting rate limits, timeouts and retries with jitter
  • Test code that depends on model calls, separating deterministic tests from evals
  • Choose between asyncio, threads and processes, and profile a slow Python program
Prerequisites
  • Basic Python: functions, classes, packages and virtual environments

Projects, Environments and Dependencies

A Python project that works on one laptop and breaks on the next is usually a dependency problem: a different version of a library got installed. Modern tooling records the exact versions in a lockfile, so every developer, test run and server installs the same thing - which turns "works on my machine" bugs into a solved problem.

Declare the project in pyproject.toml (the PEP 621 standard): name, Python version, dependencies, and dependency groups (PEP 735) for dev and test tools that should not ship. Resolve exact versions into a lockfile and install from it everywhere. uv (Astral) does all of this - creating the virtual environment, resolving, locking and running - and is much faster than pip, which matters when CI and Docker builds reinstall environments constantly.

uv init summarizer && cd summarizer      # pyproject.toml, .python-version, a starter module
uv add "pydantic>=2" httpx fastapi       # adds to [project.dependencies] and updates uv.lock
uv add --dev pytest ruff                 # goes into [dependency-groups].dev
uv sync                                  # create/refresh .venv exactly from uv.lock
uv run pytest                            # run inside the project environment

Rules that prevent most environment pain:

  • Commit the lockfile (uv.lock) for applications, and install from it in CI and Docker. Libraries publish ranges; applications pin.
  • Pin the Python version too (requires-python plus .python-version). GPU libraries such as PyTorch, vLLM and flash-attn support specific Python and CUDA combinations (Linux & the GPU Box).
  • One environment per project, never the system Python. On macOS the system python3 may be years old - a 3.9 interpreter has no asyncio.TaskGroup.
  • Python 3.14 (October 2025) makes the free-threaded (no-GIL) build officially supported but optional; most AI libraries still target the standard build.

Types and Data Validation

Type hints don't change how Python runs, but a type checker (mypy or pyright) and your editor use them to catch mistakes before they run. In LLM applications the bigger win is validating data at the boundaries - incoming requests, model outputs, tool arguments - where the data is untrusted.

from typing import Literal
from pydantic import BaseModel, Field

class Ticket(BaseModel):
    """What we expect the model to extract from a support email."""
    category: Literal["billing", "technical", "account", "other"]
    priority: int = Field(ge=1, le=4)
    summary: str = Field(max_length=300)
    order_id: str | None = None

raw = '{"category": "billing", "priority": 2, "summary": "Charged twice for March"}'
ticket = Ticket.model_validate_json(raw)     # raises ValidationError on bad or missing fields
print(ticket.priority, Ticket.model_json_schema()["properties"]["category"])

The same Pydantic model gives you three things: parsing and validation of the model's JSON output, a JSON Schema to send to structured-output APIs (Structured Outputs), and request and response validation in FastAPI (API Design for LLM Services). Fail loudly at the boundary and the rest of the code can trust its inputs.


Concurrency: Many Slow Model Calls

An LLM call spends almost all of its time waiting on the network. Calling 40 documents one after another wastes that time; asyncio lets one thread keep many calls in flight. Three things must come with it in production: a concurrency limit (providers enforce rate limits), a timeout on every call, and retries with backoff for rate-limit and transient errors.

import asyncio
import random


class RateLimited(Exception):
    """Stand-in for an HTTP 429 from a model API."""


async def call_model(prompt: str) -> str:
    """Fake model call: 0.2-0.5 s of network wait, and a 20% chance of a 429."""
    await asyncio.sleep(random.uniform(0.2, 0.5))
    if random.random() < 0.2:
        raise RateLimited(prompt)
    return f"summary of {prompt}"


async def with_retries(prompt: str, attempts: int = 5, base: float = 0.1, cap: float = 2.0) -> str:
    for attempt in range(attempts):
        try:
            async with asyncio.timeout(5):                    # never wait forever on one call
                return await call_model(prompt)
        except (RateLimited, TimeoutError):
            if attempt == attempts - 1:
                raise
            # exponential backoff with full jitter: sleep a random time up to base * 2^attempt
            await asyncio.sleep(random.uniform(0, min(cap, base * 2 ** attempt)))
    raise AssertionError("unreachable")


async def summarize_all(docs: list[str], max_concurrency: int = 8) -> list[str]:
    sem = asyncio.Semaphore(max_concurrency)                  # stay under the provider's rate limit

    async def one(doc: str) -> str:
        async with sem:
            return await with_retries(doc)

    async with asyncio.TaskGroup() as tg:                     # cancels the rest if one task fails
        tasks = [tg.create_task(one(d)) for d in docs]
    return [t.result() for t in tasks]


random.seed(0)
results = asyncio.run(summarize_all([f"doc-{i}" for i in range(40)]))
# Run with Python 3.12: 40 summaries in 2.2 s, versus ~14 s one at a time, despite simulated 429s

Why each piece is there:

  • Semaphore caps concurrent calls. Without it, gather over 10,000 documents opens 10,000 requests at once and turns the rate limit into a wall of 429s.
  • asyncio.timeout (Python 3.11+) bounds each call; a hung connection otherwise blocks a worker forever.
  • Full jitter spreads retries randomly so that many clients that failed together don't retry in lock-step and fail together again. If the API returns a Retry-After header, honour it instead.
  • TaskGroup (3.11+) gives structured concurrency: if a task fails for good, the others are cancelled and the error propagates instead of being silently lost.

asyncio, threads or processes? Use asyncio (or a thread pool) for I/O-bound work - model APIs, databases, HTTP. Use processes (multiprocessing, concurrent.futures.ProcessPoolExecutor) for CPU-bound Python, such as heavy text preprocessing, because the GIL lets only one thread run Python bytecode at a time in the standard build. GPU work releases the GIL inside the kernels, so a PyTorch training loop is limited by the GPU and data pipeline, not by threads.


Testing Code That Calls Models

Model outputs are non-deterministic and cost money, so split testing into two layers:

LayerWhat it checksHow
Unit and integration tests (pytest, every commit)Your code: prompt assembly, parsing, validation, retries, tool dispatch, error pathsReplace the model with a fake or recorded response; assert exact behaviour
Evals (scheduled, and before releases)The model plus your prompts: quality on representative inputsReal model calls, scored with metrics or judges, reported with confidence intervals (Building Your Own Evals)
import pytest
from pydantic import ValidationError

@pytest.mark.parametrize("bad", ['{"category": "refunds", "priority": 2, "summary": "x"}', "not json"])
def test_rejects_invalid_model_output(bad):
    with pytest.raises(ValidationError):
        Ticket.model_validate_json(bad)

Test the paths that matter most in production and are hardest to trigger by hand: malformed model output, timeouts, 429s, empty retrieval results, and tool errors.


Profiling

Before optimizing, measure where the time goes:

  • cProfile (standard library) - deterministic function-level profile: python -m cProfile -s cumulative app.py.
  • py-spy - a sampling profiler that attaches to a running process with negligible overhead (py-spy top --pid 1234, py-spy record -o profile.svg --pid 1234). Ideal for "the server is slow right now".
  • time.perf_counter() around suspect sections for quick checks, and tracing spans in production (Agent Observability).

In LLM applications the profile usually shows time waiting on the model, retrieval or the network, not Python itself. The fixes are then concurrency, caching and fewer or smaller model calls - not rewriting Python in another language. For GPU code, use the GPU profilers instead (CUDA Concepts & GPU Profiling).

Code quality tooling that pays for itself: ruff for linting and formatting (fast, replaces flake8, isort and black for most teams), a type checker in CI, and pre-commit to run them before each commit (Git Workflows for ML).


Check Yourself

Check yourself
0 / 5 answered
  1. Why commit uv.lock for an application?
  2. You run 5,000 model calls with asyncio.gather and get mostly HTTP 429 errors. What's the first fix?
  3. Why add random jitter to retry backoff?
  4. Which work should go to a process pool rather than asyncio or threads in standard CPython?
  5. Why test LLM application code with fake model responses rather than real calls?

Exercises

Exercise - Make a batch job production-ready

A colleague's script summarizes 20,000 tickets: results = [client.complete(t) for t in tickets]. It takes 30 hours and crashes on the first 429. List the changes you would make, in order, and estimate the new runtime if each call takes 1.5 s and the provider allows 20 concurrent requests.

Solution
  1. Make calls async (or use a thread pool) with a semaphore of about 20.
  2. Add a timeout per call and retries with exponential backoff and full jitter for 429s, 5xx and timeouts, honouring Retry-After.
  3. Validate each output with a Pydantic model; send failures to a retry queue rather than crashing.
  4. Write results incrementally (append to a file or database keyed by ticket id) so a crash resumes instead of restarting, and skip already-done ids.
  5. Log progress and failures.

Runtime: 20,000 x 1.5 s / 20 ≈ 1,500 s ≈ 25 minutes, plus retry overhead - versus 30 hours sequentially. Also check whether the provider's batch API (often cheaper, with a longer turnaround) fits the job (Cost and Latency).

Exercise - Validate model output at the boundary

Write a Pydantic model for an extraction that returns an invoice number (pattern INV- followed by 6 digits), a total (positive, at most 1,000,000), a currency (one of EUR, USD, GBP) and an optional due date. What happens when the model returns a total of -5?

Solution
from datetime import date
from typing import Literal
from pydantic import BaseModel, Field

class Invoice(BaseModel):
    number: str = Field(pattern=r"^INV-\d{6}$")
    total: float = Field(gt=0, le=1_000_000)
    currency: Literal["EUR", "USD", "GBP"]
    due_date: date | None = None

Invoice.model_validate_json(...) raises a ValidationError naming the total field. Catch it at the boundary and either retry the model call with the error message included, or route the item to human review - never let an invalid record flow downstream.

Study Notes

Must-know:

  • pyproject.toml (PEP 621) + dependency groups (PEP 735) + a committed lockfile; uv to create, lock, sync and run
  • Pin Python and dependency versions; GPU libraries constrain the Python/CUDA combination
  • Pydantic models validate untrusted data at the boundaries - requests, model outputs, tool arguments - and produce JSON Schema
  • asyncio for many I/O-bound model calls: Semaphore for rate limits, asyncio.timeout, retries with exponential backoff and full jitter, TaskGroup for structured concurrency
  • Processes for CPU-bound Python (the GIL); asyncio or threads for I/O
  • Unit tests with fake model responses on every commit; evals with real models separately
  • Profile before optimizing: cProfile, py-spy on live processes; LLM apps are usually waiting on I/O

References

Last reviewed: 2026-10

⚡AI-assisted content - always verify, always explore multiple perspectives·