20 - Capstones
Three projects that put the course together end to end, each producing something you could show an employer: a model you trained and post-trained yourself, that model served against a stated service-level objective, and a production-grade agent with evaluations and a security review. Each project lists its milestones, deliverables and a grading rubric, and builds on labs you have already run.
- Train a small language model from scratch, post-train an open base model with SFT and DPO or GRPO, and evaluate both with confidence intervals
- Serve a model against an explicit latency SLO, measure goodput under realistic load, and justify each serving choice with numbers
- Build, evaluate, secure and instrument an agent that uses MCP tools, with human approval and durable execution
- Write the reports that make engineering work credible - method, results with uncertainty, trade-offs and limitations
- Document each system's architecture with ADRs, and state the business outcome it would deliver with a cost and ROI estimate
- Capstone 1: modules 01-07, especially GPT From Scratch, SFT, DPO and GRPO with TRL and Eval Harness
- Capstone 2: modules 08-09, especially FastAPI + vLLM Endpoint
- Capstone 3: modules 13-18, especially MCP Server & Agent
- All capstones: Solutions Architecture & Communication for the architecture document and business outcome
How the Capstones Fit Together
flowchart LR
C1["๐๏ธ 1. Train & post-train<br/>a small model"] -->|"your tuned model"| C2["๐ 2. Serve it<br/>with an SLO"]
C2 -->|"your endpoint (optional)"| C3["๐ค 3. Ship a<br/>production agent"]
style C1 fill:#e8e0d4,stroke:#c8b89a
style C2 fill:#d8dfe8,stroke:#b0bac8
style C3 fill:#dde4dc,stroke:#b0c4b0
They chain - the model from capstone 1 is served in capstone 2, and that endpoint can be one of the models behind capstone 3's agent - but each stands alone: you can serve any small open model in capstone 2, or use a hosted API in capstone 3.
| # | Capstone | You deliver | Compute |
|---|---|---|---|
| 1 | Train and Post-Train a Small Model | A pretrained ~124M GPT, a post-trained 0.6B model, an evaluation report and model cards | One 24 GB GPU for about a day in total (a light track fits a laptop plus a free notebook GPU) |
| 2 | Serve It with an SLO | A deployed endpoint, load-test data, and an SLO report comparing serving configurations | One GPU for a few hours; Kubernetes optional |
| 3 | Ship a Production Agent | An MCP-backed agent with approvals, durable execution, tracing, eval suites, a threat model and a cost model | A laptop plus a local or hosted model |
Grading
Every capstone uses the same four levels per rubric criterion, and a weighted total out of 100:
| Level | Meaning |
|---|---|
| Exceeds (100%) | Everything in Meets, plus measured evidence of an improvement or insight beyond the brief |
| Meets (75%) | The deliverable is complete, correct and reproducible, and claims are backed by measurements with uncertainty |
| Partial (40%) | Present but incomplete, unmeasured, or not reproducible |
| Missing (0%) | Not delivered |
A capstone passes at 70. Reproducibility is a gate, not a criterion: if a reviewer cannot re-run your main result from the repository and README, the capstone is returned before grading.
What Every Submission Includes
- A public or shareable repository with pinned dependencies, a README with exact commands, and seeds.
- A report (2-5 pages) - question, method, results with confidence intervals or error bars, what didn't work, and limitations.
- The artifacts each capstone lists (model cards, dashboards, traces, threat model).
- An architecture document (1-3 pages): context and container diagrams, the non-functional requirements you designed for, and ADRs for the two or three decisions that were hardest to reverse - each with its evidence and a trigger to revisit it (Architecture Docs & ADRs).
- A business outcome: who would use this, the metric it moves, and a unit cost and ROI estimate with a sensitivity table - stating which savings are cash and which are capacity (Use-Case Qualification & ROI). For a learning project, a realistic hypothetical customer is fine; say so.
- A 5-minute walkthrough (recording or live) of the main result.
Together, the three capstones are the "three end-to-end projects" of the Readiness Self-Assessment - code, metrics, architecture, trade-offs and business outcome.
Honest negative results are fine and often the most instructive part of a report - the labs in this course report several. Unmeasured claims are not.
Previous: 19 - Solutions Architecture & Communication ยท Next: Knowledge Check
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.