Docker for GPU Inference
The One-Line Definition
A GPU inference image is a normal Python image plus CUDA user-space libraries (from pip wheels or a CUDA base image), run with a container runtime that injects the host's GPU driver - and, when kernels must be compiled, a multi-stage build so the compiler toolchain never ships to production.
- Explain which layer supplies the GPU driver, the CUDA libraries and the framework in a GPU container
- Diagnose
torch.cuda.is_available() == Falsein a container by checking GPU access, wheel variant and driver version in order - Choose between a slim Python base with CUDA wheels, a CUDA base image and a vendor image such as
vllm/vllm-openai - Write a multi-stage Dockerfile that compiles in a
-develstage and ships a-runtimestage, with weights mounted rather than copied
- FastAPI + vLLM Endpoint - the service being containerized
Packaging a model server into a container is not fundamentally different from packaging any other web service - except the container needs to talk to a physical GPU on the host machine, and the base image alone can be several gigabytes before a single line of application code is added. This page is about what actually changes versus a normal web-app Dockerfile, and how to keep the image from becoming unnecessarily bloated.
This page assumes the FastAPI + vLLM serving stack from Inference & Serving as the thing being containerized, and builds toward the Helm Chart & Release Checklist Code Lab where this image gets deployed to Kubernetes.
flowchart LR
Base["๐งฑ Base image\n(python-slim + CUDA wheels,\nnvidia/cuda, or vendor image)"] --> Build["๐ง Build stage\ncompile deps, download\nmodel weights/cache"]
Build --> Runtime["๐ Runtime stage\nminimal CUDA runtime\n+ app code only"]
Runtime --> Registry["๐ฆ Push to registry\n(tagged, scanned)"]
style Base fill:#d8dfe8,stroke:#b0bac8
style Build fill:#e8e0d4,stroke:#c8b89a
style Runtime fill:#dde4dc,stroke:#b0c4b0
style Registry fill:#ddd8e4,stroke:#b8b0c8
What a GPU Container Actually Needs
A GPU program in a container needs three layers, and only the last one comes from your image:
- The GPU driver - on the host, never in the image. The driver (including
libcuda.so) is tied to the host's kernel and GPU. The NVIDIA Container Toolkit injects the driver libraries and device nodes into the container when you run it with GPU access -docker run --gpus all, or in Kubernetes the NVIDIA device plugin plus the NVIDIA runtime (runtimeClassName: nvidiawhere required). - CUDA user-space libraries (runtime, cuBLAS, cuDNN, NCCL) - these can come from the image, and they must be supported by the host driver version (newer CUDA needs a newer driver).
- Your framework - PyTorch, vLLM and similar.
So where do the CUDA libraries come from? Pip wheels bring their own: the Linux CUDA builds of torch and vllm depend on nvidia-* pip packages containing the CUDA runtime, cuBLAS, cuDNN and NCCL. A plain Python base image therefore works for many inference servers:
# Works for pip-installed CUDA wheels, as long as the container runs with --gpus all
FROM python:3.12-slim
RUN pip install --no-cache-dir vllm
# docker run --gpus all ... -> torch.cuda.is_available() == True
A CUDA base image (nvidia/cuda:<version>-runtime-<os>, or a vendor image such as vllm/vllm-openai that pins CUDA, PyTorch and vLLM together) is the better choice when:
- you compile CUDA code (flash-attn or custom kernels from source) - that needs the
-develimage'snvccand headers, typically in a build stage; - a library expects system CUDA libraries rather than pip-provided ones (some TensorRT and custom C++ builds);
- you want one known-good combination of CUDA, cuDNN and NCCL across every service.
# CUDA runtime base + a virtualenv
FROM nvidia/cuda:12.8.1-runtime-ubuntu24.04
RUN apt-get update && apt-get install -y python3.12 python3.12-venv
# Ubuntu 24.04 blocks system-wide pip installs (PEP 668), so install into a virtualenv
RUN python3.12 -m venv /opt/venv && /opt/venv/bin/pip install vllm
ENV PATH="/opt/venv/bin:$PATH"
When torch.cuda.is_available() is False in a container, check in this order: the container was started without GPU access (no --gpus, missing device plugin or runtime class); a CPU-only wheel was installed (for example from the CPU index URL); or the host driver is too old for the wheel's CUDA version (nvidia-smi shows the maximum CUDA version the driver supports).
Image Size Considerations
CUDA base images are large before any application code is added - often 3-8 GB just for the base layer, versus a few hundred MB for a plain Python image. That size matters operationally: bigger images take longer to pull on every new pod/node, slow down autoscaling and rolling deploys, and cost more in registry storage and network egress at scale.
Where the size actually comes from, and what to do about each source:
| Source | Typical size | Mitigation |
|---|---|---|
| CUDA runtime base image | 2-4 GB | Use -runtime variant, not -devel (no compiler toolchain needed at runtime) |
| PyTorch + CUDA-linked wheels | 3-6 GB | Pin exact versions; avoid pulling both CPU and GPU wheel variants. The wheels carry their own CUDA libraries, so on a CUDA base image you ship two copies - one reason the framework's own image or a slim base is often smaller |
| Model weights baked into the image | Multi-GB per model | Prefer mounting weights from a volume/object store at startup over COPY-ing them into the image (see below) |
pip/apt build caches, .git, test files | 100s of MB - GBs if unmanaged | --no-cache-dir on pip, rm -rf /var/lib/apt/lists/*, .dockerignore |
Should model weights live inside the image or be mounted at runtime? Baking multi-GB weights into the image makes every rebuild re-push the full weight layer even for a one-line code change, and couples model versioning to container versioning. Mounting weights from a PersistentVolume, object storage (S3/GCS), or the HuggingFace Hub cache at container start keeps the image itself small and lets model version and code version roll independently - this is the pattern used in the Helm Chart Code Lab's values.yaml (modelPath mounted as a volume rather than COPY'd).
Multi-Stage Builds
A multi-stage build compiles and installs everything needed to build the application in one throwaway stage, then copies only the finished artifacts into a clean, minimal final image - the compiler, build tools, and intermediate files never make it into what actually ships and runs in production.
# ---- Stage 1: build ----
FROM nvidia/cuda:12.8.1-devel-ubuntu24.04 AS builder
RUN apt-get update && apt-get install -y python3.12 python3-pip python3.12-venv
WORKDIR /build
COPY requirements.txt .
RUN python3.12 -m venv /opt/venv \
&& /opt/venv/bin/pip install --no-cache-dir -r requirements.txt
# ---- Stage 2: runtime ----
FROM nvidia/cuda:12.8.1-runtime-ubuntu24.04
RUN apt-get update && apt-get install -y python3.12 --no-install-recommends \
&& rm -rf /var/lib/apt/lists/*
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
WORKDIR /app
COPY serve.py .
EXPOSE 8000
CMD ["python3.12", "serve.py"]
Why this matters here specifically: the -devel CUDA image (with nvcc, headers, and build tools for compiling packages like flash-attn from source) can be 2-3 GB larger than the -runtime variant. Building in a -devel stage and shipping only the resulting virtualenv on a -runtime base can meaningfully cut final image size while keeping the exact same compiled binaries. See the Code Lab Dockerfile for a complete, runnable version of this pattern.
Study Notes
Must-know for interviews:
- The host provides the driver (injected by the NVIDIA Container Toolkit when the container runs with GPU access); the image provides CUDA user-space libraries - pip CUDA wheels bundle them, CUDA base images provide them system-wide and are needed for compiling kernels
- CUDA base images are large (multi-GB) before any app code - use
-runtimenot-develvariants for the final stage, and avoid baking model weights into the image - Multi-stage builds let you compile in a fat
-develstage and ship only the resulting artifacts on a slim-runtimestage, cutting final image size without changing the compiled output - Model weights are generally better mounted at container start (volume/object store) than
COPY'd into the image, so model version and code version can roll independently
Check Yourself
- Where does
libcuda.so(the GPU driver library) come from inside a running GPU container? - Which situation genuinely requires a CUDA
-develimage (in at least one build stage)? - Why does
torch.cuda.is_available()returnFalseinside a container even on a GPU host? - What's the practical difference between a CUDA
-develand-runtimeimage tag? - Why should model weights usually not be
COPY'd into the Docker image? - A teammate suggests
FROM python:3.12-slimand installing CUDAtorchvia pip, arguing it's simpler than a CUDA base image. Is that wrong?
Exercises
A container built FROM python:3.12-slim with pip install torch prints torch.cuda.is_available() == False on a host with an H100. List the checks you would run, in order, and the command or evidence for each.
Hint
Three causes cover almost every case.
Solution
- GPU access - was it started with
docker run --gpus all(or, in Kubernetes, does the pod requestnvidia.com/gpuon a node with the device plugin and NVIDIA runtime)? Inside the container,nvidia-smiorls /dev/nvidia*should succeed. - Wheel variant -
python -c "import torch; print(torch.version.cuda)".Nonemeans a CPU-only wheel was installed; reinstall from the CUDA index. - Driver version -
nvidia-smion the host shows the highest CUDA version the driver supports; it must be at least the wheel'storch.version.cuda(minor-version compatibility aside). Upgrade the driver or pick a wheel built for an older CUDA.
An image is 14 GB: nvidia/cuda:*-devel base (โ7 GB), CUDA torch + vLLM wheels (โ5 GB), and a 2 GB model copied in with COPY. Propose changes and estimate the new size.
Solution
- Move compilation (if any) to a
-develbuild stage and ship on a slim or-runtimebase: the-develtoolkit is several GB you do not need at run time. Since the wheels bundle their own CUDA libraries, apython:3.12-slimbase (โ0.15 GB) is enough if nothing is compiled. - Mount the weights from a volume or object store instead of
COPY- removes 2 GB and decouples model and code versions. --no-cache-dir, cleaned apt lists and a.dockerignore.
Result: roughly 5-6 GB (wheels plus a slim base), down from 14 GB - and a code change no longer re-pushes the weights.
References
- NVIDIA, Container Toolkit documentation (2026)
- NVIDIA, CUDA compatibility - driver vs toolkit versions (2026)
- PyTorch, Get Started - install matrix (2026)
- Docker, Multi-stage builds (2026)
Last reviewed: 2026-09