Readiness Self-Assessment
A self-assessment for GenAI engineering roles: score yourself 0-3 on fourteen domains, compare with the target for each, and use the links to the modules and labs that close the gap. Retake it after each module or capstone - the scores are only useful if they are based on evidence you could show someone.
- Score yourself 0-3 on each of the 14 domains using the evidence descriptions, not a feeling
- Identify every core domain below 2 and the modules and labs that address it
- Plan three end-to-end projects that show code, metrics, architecture, trade-offs and business outcome
- None - take it before starting the course and again after each track
The Scale
| Score | Meaning |
|---|---|
| 0 | Cannot explain or build it |
| 1 | Recognize the concepts; can follow a tutorial |
| 2 | Can build it independently |
| 3 | Can design it, debug it and explain the trade-offs |
Readiness threshold: no score below 2 in core areas, and at least three end-to-end projects where you can show the code, metrics, architecture, trade-offs and business outcome.
Domains, Targets and Where to Learn Them
Coverage says how fully this course covers the domain: Full means the notes and labs can take you to the target. Every domain is now covered in full; the links point to the notes and labs that build each one.
What a 2 and a 3 Look Like
Score yourself on what you have actually done, and could show.
| # | You're at 2 when you can... | You're at 3 when you can also... |
|---|---|---|
| 1 | Build a typed, tested Python service with a REST or streaming API, containerize it, and work on a remote Linux GPU machine with Git branches and PRs | Design the API contract (streaming, idempotency, rate limits, versioning), and debug a container, driver or dependency problem from first principles |
| 2 | Deploy a model server on Kubernetes with Helm, GPU scheduling and least-privilege IAM | Design a private, multi-environment deployment (networking, identity, IaC, autoscaling signal) and defend it against alternatives on another cloud |
| 3 | Work through attention shapes, softmax and cross-entropy by hand, and compute a confidence interval for an eval score | Explain optimizer behaviour, low-rank adaptation and training instabilities mathematically, and choose the right statistical test for a comparison |
| 4 | Write a training loop from scratch and fine-tune a Hugging Face model with your own data | Diagnose a training run that won't converge, or a memory blow-up, from the code and the curves |
| 5 | Implement a small GPT and explain each component, and how a tokenizer turns text into ids | Explain why an architecture choice (GQA, MoE, RoPE scaling, vocabulary size) changes quality, memory or cost, with numbers |
| 6 | Run LoRA/QLoRA and preference tuning on curated data, and show a measured improvement over the base model | Decide between prompting, RAG, SFT, DPO and RL for a problem, and design the data pipeline and checks for it |
| 7 | Build a task-specific eval set, use a validated LLM judge, and run a red-team scan reporting ASR and over-refusal | Design an evaluation and safety programme for a product - gates, online evals, attacker budgets, grader validation - and explain what each number can't tell you |
| 8 | Estimate the GPU memory of a training or serving job and fix an out-of-memory error | Profile a slow step, identify whether it is compute-, memory- or communication-bound, and fix it |
| 9 | Run a multi-GPU training job with DDP or FSDP | Choose a parallelism strategy (data, tensor, pipeline, ZeRO stage) for a model and cluster, and estimate the communication cost |
| 10 | Serve a model with vLLM behind an API and load-test it | Tune batching, KV cache, precision and parallelism against an SLO, and explain the goodput trade-offs |
| 11 | Build and evaluate a RAG pipeline with hybrid retrieval and reranking | Design retrieval over enterprise sources with permissions, freshness and governance, and measure and fix retrieval failures |
| 12 | Build a tool-using agent with an MCP server, bounds and human approval | Design a production agent's security model (injection, trifecta, credentials), durability and evaluation, and choose workflow vs agent with evidence |
| 13 | Instrument a service with traces and metrics and set an SLO with alerts | Run an incident and a blameless postmortem, design burn-rate alerting and cost controls, and catch silent quality regressions |
| 14 | Run a discovery conversation and write a design doc with diagrams and ADRs | Qualify a use case with an ROI model, present it to executives and engineers, and plan a pilot with exit criteria to production |
Three End-to-End Projects
The Capstones are designed as the three projects - each asks for code, metrics with uncertainty, an architecture document with ADRs, trade-offs and a business outcome.
| Capstone | Domains it gives evidence for |
|---|---|
| 1 - Train and Post-Train a Small Model | 3, 4, 5, 6, 7, 8, 9 |
| 2 - Serve It with an SLO | 1, 2, 8, 10, 13 |
| 3 - Ship a Production Agent | 7, 11, 12, 13 |
| All three | 14 - every capstone requires a design document, ADRs and a business outcome |
If a domain you need is not covered by these, build a fourth project around it - for example, a permission-aware RAG system over a real document set for domain 11.
How to Use It
- Score all fourteen domains honestly, using the evidence columns above.
- List every core domain below 2. Those come first, before polishing domains already at 2.
- For each, work through the linked notes, then do the lab - reading moves you from 0 to 1; building moves you to 2.
- Move to 3 by explaining trade-offs to someone else: the module Q&A banks and the system designs are practice for that.
- Rescore after each track, and keep the evidence (repositories, reports, dashboards) - it is what you will show in interviews.