These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Check yourself 0 / 4 answered
What distinguishes orchestrator-workers from parallelization by sectioning?
A In orchestrator-workers the sub-tasks are decided at run time by a model; in sectioning they are fixed in advance B Orchestrator-workers always uses more models C Sectioning can't run in parallel D There is no difference When is an evaluator-optimizer loop most likely to improve results?
A When the evaluator has information the generator lacked, such as test results or a rubric B Whenever the same model re-reads its answer C When the loop runs at least ten rounds D Only with a larger evaluator model A router sends 5% of refund requests to the general-FAQ handler. How would you find and fix this?
Show answer Best-of-5 with the same model choosing its favourite gives little gain; best-of-5 selected by unit tests gives a large one. Why?
Show answer
Check yourself 0 / 4 answered
Why is model self-reported confidence a poor basis for a confidence gate?
A It is poorly calibrated; external signals such as verifiers or tests are more reliable B Models can't output numbers C It is too slow D It is always too low What does using a sub-agent as a tool mainly buy you?
A Context isolation - exploration noise stays out of the main agent's context B Lower total token cost C Guaranteed correctness D Peer-to-peer negotiation between agents Which verification loop is most likely to fix a bug in generated code?
A Running the tests and feeding the failures back B Asking the model to re-read its code C Asking a second model whether the code looks right D Increasing temperature and regenerating What should a handoff to a human agent contain?
Show answer
Check yourself 0 / 5 answered
According to Anthropic's analysis on BrowseComp, what explained most of the performance variance?
A Token usage B The number of agents C Model temperature D Prompt length Which task is the worst fit for parallel sub-agents?
A A sequential refactoring where each step depends on the previous change B Researching the pricing of ten competitors C Summarising twenty independent documents D Checking five hypotheses against different data sources What is the main difference between handoffs and orchestrator-subagents?
A Handoffs transfer control to one active agent at a time; an orchestrator keeps control and delegates sub-tasks B Handoffs need more tokens C Orchestrators can't run agents in parallel D Handoffs require A2A Which MAST category does 'the system stopped before verifying the result' belong to, and what mitigates it?
Show answer Why might a multi-agent debate beat a single agent in a paper yet not in your system?
Show answer
Check yourself 0 / 4 answered
What problem does an artifact store (sub-agents write outputs and return references) solve?
A Information loss and token overhead from passing large outputs through the lead agent B Authentication between agents C Model selection D Prompt injection Why keep a structured task ledger in the orchestrator's state?
A It survives compaction, lets the orchestrator detect stalls and re-plan, and gives humans a readable view B It replaces tracing C It reduces the number of sub-agents D Frameworks require it What is a rainbow deployment for long-running agents?
A Running old and new versions side by side and shifting traffic gradually so in-flight runs finish on their original version B Deploying each agent in a different colour-coded environment C Randomly routing each step to a different model D Canary-testing prompts on 1% of tokens Your orchestrator spawns 12 sub-agents for 'What is the capital of Australia?'. What do you change?
Show answer
Check yourself 0 / 4 answered
Why do strategies use the docstring examples but grading uses HumanEval's hidden tests?
A So strategies aren't scored on the same checks they optimised against B The hidden tests are faster C The docstring examples are wrong D To use fewer tokens Why report a paired difference rather than two separate pass rates?
A Both strategies ran on the same problems, so pairing removes between-problem variance and gives a much tighter interval B Paired CIs are always positive C It is required by HumanEval D It hides ceiling effects self_refine scored lower than single. Does that prove self-review hurts?
Show answer Why did test_feedback cost only 5% more model calls than single?
Show answer