Python for AI Engineering
Most AI engineering is ordinary software engineering in Python: managing environments and dependencies, writing typed and tested code, and handling many slow network calls to models at once. This note covers the parts that matter most for LLM work - project setup with uv and pyproject.toml, type hints and Pydantic models, asyncio for concurrent model calls with rate limits, timeouts and retries, testing code that calls non-deterministic models, and profiling.
- Set up a reproducible Python project with pyproject.toml, a lockfile and dependency groups
- Use type hints and Pydantic models to validate data at the boundaries of an LLM application
- Run many model calls concurrently with asyncio while respecting rate limits, timeouts and retries with jitter
- Test code that depends on model calls, separating deterministic tests from evals
- Choose between asyncio, threads and processes, and profile a slow Python program
- Basic Python: functions, classes, packages and virtual environments
Projects, Environments and Dependencies
A Python project that works on one laptop and breaks on the next is usually a dependency problem: a different version of a library got installed. Modern tooling records the exact versions in a lockfile, so every developer, test run and server installs the same thing - which turns "works on my machine" bugs into a solved problem.
Declare the project in pyproject.toml (the PEP 621 standard): name, Python version, dependencies, and dependency groups (PEP 735) for dev and test tools that should not ship. Resolve exact versions into a lockfile and install from it everywhere. uv (Astral) does all of this - creating the virtual environment, resolving, locking and running - and is much faster than pip, which matters when CI and Docker builds reinstall environments constantly.
uv init summarizer && cd summarizer # pyproject.toml, .python-version, a starter module
uv add "pydantic>=2" httpx fastapi # adds to [project.dependencies] and updates uv.lock
uv add --dev pytest ruff # goes into [dependency-groups].dev
uv sync # create/refresh .venv exactly from uv.lock
uv run pytest # run inside the project environment
Rules that prevent most environment pain:
- Commit the lockfile (
uv.lock) for applications, and install from it in CI and Docker. Libraries publish ranges; applications pin. - Pin the Python version too (
requires-pythonplus.python-version). GPU libraries such as PyTorch, vLLM and flash-attn support specific Python and CUDA combinations (Linux & the GPU Box). - One environment per project, never the system Python. On macOS the system
python3may be years old - a 3.9 interpreter has noasyncio.TaskGroup. - Python 3.14 (October 2025) makes the free-threaded (no-GIL) build officially supported but optional; most AI libraries still target the standard build.
Types and Data Validation
Type hints don't change how Python runs, but a type checker (mypy or pyright) and your editor use them to catch mistakes before they run. In LLM applications the bigger win is validating data at the boundaries - incoming requests, model outputs, tool arguments - where the data is untrusted.
from typing import Literal
from pydantic import BaseModel, Field
class Ticket(BaseModel):
"""What we expect the model to extract from a support email."""
category: Literal["billing", "technical", "account", "other"]
priority: int = Field(ge=1, le=4)
summary: str = Field(max_length=300)
order_id: str | None = None
raw = '{"category": "billing", "priority": 2, "summary": "Charged twice for March"}'
ticket = Ticket.model_validate_json(raw) # raises ValidationError on bad or missing fields
print(ticket.priority, Ticket.model_json_schema()["properties"]["category"])
The same Pydantic model gives you three things: parsing and validation of the model's JSON output, a JSON Schema to send to structured-output APIs (Structured Outputs), and request and response validation in FastAPI (API Design for LLM Services). Fail loudly at the boundary and the rest of the code can trust its inputs.
Concurrency: Many Slow Model Calls
An LLM call spends almost all of its time waiting on the network. Calling 40 documents one after another wastes that time; asyncio lets one thread keep many calls in flight. Three things must come with it in production: a concurrency limit (providers enforce rate limits), a timeout on every call, and retries with backoff for rate-limit and transient errors.
import asyncio
import random
class RateLimited(Exception):
"""Stand-in for an HTTP 429 from a model API."""
async def call_model(prompt: str) -> str:
"""Fake model call: 0.2-0.5 s of network wait, and a 20% chance of a 429."""
await asyncio.sleep(random.uniform(0.2, 0.5))
if random.random() < 0.2:
raise RateLimited(prompt)
return f"summary of {prompt}"
async def with_retries(prompt: str, attempts: int = 5, base: float = 0.1, cap: float = 2.0) -> str:
for attempt in range(attempts):
try:
async with asyncio.timeout(5): # never wait forever on one call
return await call_model(prompt)
except (RateLimited, TimeoutError):
if attempt == attempts - 1:
raise
# exponential backoff with full jitter: sleep a random time up to base * 2^attempt
await asyncio.sleep(random.uniform(0, min(cap, base * 2 ** attempt)))
raise AssertionError("unreachable")
async def summarize_all(docs: list[str], max_concurrency: int = 8) -> list[str]:
sem = asyncio.Semaphore(max_concurrency) # stay under the provider's rate limit
async def one(doc: str) -> str:
async with sem:
return await with_retries(doc)
async with asyncio.TaskGroup() as tg: # cancels the rest if one task fails
tasks = [tg.create_task(one(d)) for d in docs]
return [t.result() for t in tasks]
random.seed(0)
results = asyncio.run(summarize_all([f"doc-{i}" for i in range(40)]))
# Run with Python 3.12: 40 summaries in 2.2 s, versus ~14 s one at a time, despite simulated 429s
Why each piece is there:
Semaphorecaps concurrent calls. Without it,gatherover 10,000 documents opens 10,000 requests at once and turns the rate limit into a wall of 429s.asyncio.timeout(Python 3.11+) bounds each call; a hung connection otherwise blocks a worker forever.- Full jitter spreads retries randomly so that many clients that failed together don't retry in lock-step and fail together again. If the API returns a
Retry-Afterheader, honour it instead. TaskGroup(3.11+) gives structured concurrency: if a task fails for good, the others are cancelled and the error propagates instead of being silently lost.
asyncio, threads or processes? Use asyncio (or a thread pool) for I/O-bound work - model APIs, databases, HTTP. Use processes (multiprocessing, concurrent.futures.ProcessPoolExecutor) for CPU-bound Python, such as heavy text preprocessing, because the GIL lets only one thread run Python bytecode at a time in the standard build. GPU work releases the GIL inside the kernels, so a PyTorch training loop is limited by the GPU and data pipeline, not by threads.
Testing Code That Calls Models
Model outputs are non-deterministic and cost money, so split testing into two layers:
| Layer | What it checks | How |
|---|---|---|
| Unit and integration tests (pytest, every commit) | Your code: prompt assembly, parsing, validation, retries, tool dispatch, error paths | Replace the model with a fake or recorded response; assert exact behaviour |
| Evals (scheduled, and before releases) | The model plus your prompts: quality on representative inputs | Real model calls, scored with metrics or judges, reported with confidence intervals (Building Your Own Evals) |
import pytest
from pydantic import ValidationError
@pytest.mark.parametrize("bad", ['{"category": "refunds", "priority": 2, "summary": "x"}', "not json"])
def test_rejects_invalid_model_output(bad):
with pytest.raises(ValidationError):
Ticket.model_validate_json(bad)
Test the paths that matter most in production and are hardest to trigger by hand: malformed model output, timeouts, 429s, empty retrieval results, and tool errors.
Profiling
Before optimizing, measure where the time goes:
cProfile(standard library) - deterministic function-level profile:python -m cProfile -s cumulative app.py.py-spy- a sampling profiler that attaches to a running process with negligible overhead (py-spy top --pid 1234,py-spy record -o profile.svg --pid 1234). Ideal for "the server is slow right now".time.perf_counter()around suspect sections for quick checks, and tracing spans in production (Agent Observability).
In LLM applications the profile usually shows time waiting on the model, retrieval or the network, not Python itself. The fixes are then concurrency, caching and fewer or smaller model calls - not rewriting Python in another language. For GPU code, use the GPU profilers instead (CUDA Concepts & GPU Profiling).
Code quality tooling that pays for itself: ruff for linting and formatting (fast, replaces flake8, isort and black for most teams), a type checker in CI, and pre-commit to run them before each commit (Git Workflows for ML).
Check Yourself
- Why commit uv.lock for an application?
- You run 5,000 model calls with asyncio.gather and get mostly HTTP 429 errors. What's the first fix?
- Why add random jitter to retry backoff?
- Which work should go to a process pool rather than asyncio or threads in standard CPython?
- Why test LLM application code with fake model responses rather than real calls?
Exercises
A colleague's script summarizes 20,000 tickets: results = [client.complete(t) for t in tickets]. It takes 30 hours and crashes on the first 429. List the changes you would make, in order, and estimate the new runtime if each call takes 1.5 s and the provider allows 20 concurrent requests.
Solution
- Make calls async (or use a thread pool) with a semaphore of about 20.
- Add a timeout per call and retries with exponential backoff and full jitter for 429s, 5xx and timeouts, honouring
Retry-After. - Validate each output with a Pydantic model; send failures to a retry queue rather than crashing.
- Write results incrementally (append to a file or database keyed by ticket id) so a crash resumes instead of restarting, and skip already-done ids.
- Log progress and failures.
Runtime: 20,000 x 1.5 s / 20 ≈ 1,500 s ≈ 25 minutes, plus retry overhead - versus 30 hours sequentially. Also check whether the provider's batch API (often cheaper, with a longer turnaround) fits the job (Cost and Latency).
Write a Pydantic model for an extraction that returns an invoice number (pattern INV- followed by 6 digits), a total (positive, at most 1,000,000), a currency (one of EUR, USD, GBP) and an optional due date. What happens when the model returns a total of -5?
Solution
from datetime import date
from typing import Literal
from pydantic import BaseModel, Field
class Invoice(BaseModel):
number: str = Field(pattern=r"^INV-\d{6}$")
total: float = Field(gt=0, le=1_000_000)
currency: Literal["EUR", "USD", "GBP"]
due_date: date | None = None
Invoice.model_validate_json(...) raises a ValidationError naming the total field. Catch it at the boundary and either retry the model call with the error message included, or route the item to human review - never let an invalid record flow downstream.
Study Notes
Must-know:
- pyproject.toml (PEP 621) + dependency groups (PEP 735) + a committed lockfile;
uvto create, lock, sync and run - Pin Python and dependency versions; GPU libraries constrain the Python/CUDA combination
- Pydantic models validate untrusted data at the boundaries - requests, model outputs, tool arguments - and produce JSON Schema
- asyncio for many I/O-bound model calls: Semaphore for rate limits,
asyncio.timeout, retries with exponential backoff and full jitter, TaskGroup for structured concurrency - Processes for CPU-bound Python (the GIL); asyncio or threads for I/O
- Unit tests with fake model responses on every commit; evals with real models separately
- Profile before optimizing: cProfile, py-spy on live processes; LLM apps are usually waiting on I/O
References
- Python, What's New in Python 3.14 (2025); PEP 779 - Criteria for supported status for free-threaded Python (2025)
- PEP 735 - Dependency Groups in pyproject.toml (2024); PEP 751 - A file format to record Python dependencies for installation reproducibility (2025)
- Astral, uv documentation and Ruff (2026)
- Pydantic documentation (2026); pytest documentation (2026)
- Brooker, Exponential Backoff and Jitter (AWS Architecture Blog, 2015)
- Frederickson, py-spy: Sampling profiler for Python programs (2018-2026)
Last reviewed: 2026-10