Model Lifecycle and Rollout
Deploying a new model version isn't a light switch - it's a controlled rollout with an exit plan. The best teams decide in advance what "worse" looks like and how to back out, before a single real user sees the new version. Deciding that mid-incident, under pressure, is how bad rollbacks happen.
Model lifecycle management covers four concerns: tracking which model version is running where (registry), how a new version is introduced without a hard cutover (rollout strategy), how the new version is compared against the old one on live traffic (champion-challenger), and what automatically triggers a rollback (pre-decided criteria, not judgment calls made during an incident).
- Explain what a model registry stores and how aliases (such as
@champion) drive promotion - Compare blue/green, canary, shadow and champion-challenger rollouts and pick one for a scenario
- Write quantified rollback criteria before launch, covering quality, latency, errors and cost
- Kubernetes & Helm
- Evaluation & Benchmarks - offline evaluation
Model Registry and Versioning
A model registry is a paper trail - which model version is currently serving, what data it was trained on, who approved it, and what the previous version was in case you need to go back.
Tools like MLflow Model Registry track a model's lineage: training run, metrics, artifacts, and which version is currently promoted. Current MLflow does this with aliases (mutable named pointers such as @champion and @challenger) and tags; the older fixed stages (Staging / Production / Archived) have been deprecated since MLflow 2.9. Serving code loads models:/support-bot@champion, and promotion is just moving the alias. This is metadata infrastructure - it doesn't move traffic by itself; it's the source of truth that rollout tooling (Kubernetes, a feature-flag/traffic-splitting layer) reads from when deciding what to deploy where. See Fine-Tuning Lab for how the base-vs-tuned benchmark table this course builds earlier becomes the registry entry that justifies moving the @champion alias to the new version.
Rollout Strategies
Instead of switching every user to the new model at once, roll it out gradually - a small slice of traffic first, then more, watching for problems the whole way. If something looks wrong, you've only affected a fraction of users, and you can revert instantly.
- Blue/green - two full environments running simultaneously (old = blue, new = green); traffic cuts over atomically at the load balancer/Service level, and rollback is just cutting back
- Canary - a small percentage of traffic (e.g. 5%) routed to the new version while most stays on the old one; the percentage ramps up only if health/quality metrics hold
- Shadow - the new version receives a mirrored copy of live requests and its responses are logged but never returned to users; no user risk, but double the compute and no signal on user reactions
- Champion-challenger - both versions run continuously on live traffic (often via shadow traffic or a persistent small split), with the challenger promoted to champion only after it wins on the metrics that matter over a meaningful sample size
flowchart LR
T["๐ฆ Incoming Traffic\n100%"] --> S{"Split"}
S -->|"95%"| C["๐ Champion\n(current model)"]
S -->|"5%"| Ch["๐ฅ Challenger\n(new model)"]
C --> M["๐ Metrics Compare"]
Ch --> M
M -->|"challenger wins"| P["โ
Promote & ramp"]
M -->|"challenger loses"| R["โฉ๏ธ Roll back to 0%"]
style C fill:#d8dfe8,stroke:#b0bac8
style Ch fill:#e8e0d4,stroke:#c8b89a
style M fill:#dde4dc,stroke:#b0c4b0
Rollback Criteria, Decided Up Front
The single biggest cause of bad rollbacks isn't the rollback itself - it's not having agreed in advance what "bad enough to roll back" actually means. Write the thresholds down before launch, so the decision during an incident is "check the number against the line we already drew," not a debate.
Rollback criteria should be quantitative and pre-agreed: error rate above X%, p95 latency above Y ms, eval/quality score below Z, or a spike in a specific failure mode (refusal rate, tool-call failure rate - see Agent Evaluation and Benchmarks). Automating the rollback trigger (a monitor that reverts the traffic split when a threshold breaches) removes the temptation to "wait and see" during an incident.
Study Notes
Must-know for interviews:
- A model registry (e.g. MLflow) is metadata/lineage tracking, not a deployment mechanism by itself
- Blue/green cuts over atomically; canary ramps gradually; champion-challenger compares continuously on live traffic
- Rollback criteria must be decided and quantified before launch, not during an incident
- The base-vs-tuned benchmark table from a fine-tuning workflow is exactly the evidence a registry promotion decision needs
Check Yourself
- You want to see how a new model behaves on real traffic without any user seeing its answers. Which strategy fits?
- In MLflow's current registry, how is "the version production should serve" usually marked?
- What's the difference between canary and champion-challenger?
- Why is "we'll decide if it's bad enough to roll back when we see it" a risky plan?
- What does a model registry actually store?
Exercises
A fine-tuned support model replaces the current one through a canary (5% โ 25% โ 100%). Write four quantified rollback criteria and say how each is measured.
Solution
Example (numbers depend on the product):
| Criterion | Threshold | Measured by |
|---|---|---|
| Quality | Resolution rate or judge score more than 2 points below the champion, significant at 95% | Online eval on a sample of canary vs control traffic |
| Latency | p95 TTFT above 800 ms for 10 minutes | Serving metrics per version label |
| Errors | 5xx or malformed-output rate above 1% | Gateway logs and output validators |
| Cost | Tokens per resolved conversation up more than 20% | Usage logs per version |
Each stage runs long enough to collect the samples the quality test needs; any breach reverts the traffic split (or moves the alias back) automatically.
References
- MLflow, Model Registry - aliases (2026)
- Google, The Site Reliability Workbook - Canarying Releases (2018)
Last reviewed: 2026-09