Quiz · 18 · Agent Engineering
66 questions from 10 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Harness Engineering
Check yourself
0 / 4 answered
- Two products use the same model; one completes far more long coding tasks. What most plausibly explains it?
- In the initializer/worker design, what does each worker session do first?
- Why put verification in a stop hook rather than in the system prompt?
- How do you decide whether a harness component is still needed after a model upgrade?
Loops and Graphs
Check yourself
0 / 4 answered
- A task happens twice a year and takes an hour by hand. Building a reliable loop would take two days. Build it?
- A loop wraps a model that critiques and revises its own answer three times. What limits its improvement?
- Which situation most clearly calls for a checkpointed graph rather than a loop?
- Why must graph nodes be idempotent?
Context Engineering for Agents
Check yourself
0 / 4 answered
- A model has a 1M-token window. Why still curate an agent's context?
- Which technique keeps full fidelity of a 2 MB log while keeping it out of the context?
- Why do coding agents usually explore a repository with search and file reads instead of loading it into the prompt?
- How does compaction interact with prompt caching?
Coding Agents
Check yourself
0 / 4 answered
- What is AGENTS.md?
- A team's agent-written PRs often break conventions that reviewers catch days later. What's the most effective fix?
- What did METR's 2025 randomised trial find?
- When is a background (cloud) coding agent a better fit than an interactive one?
Computer-Use and Browser Agents
Check yourself
0 / 4 answered
- A task needs data from a SaaS tool that has a REST API and a web UI. What should the agent use?
- Why is a browser agent with the user's logged-in session especially risky?
- What is 'takeover' in a computer-use harness?
- Why does reliability rather than peak capability limit unattended deployment?
Skills and Memory
Check yourself
0 / 4 answered
- Which SKILL.md frontmatter fields are required?
- An agent never uses an installed skill even when it should. What is the first thing to fix?
- Where should the command to run a repository's tests go?
- Why is 'store everything the user's emails say about their preferences' a risky memory write policy?
Verifiers, Environments and Agent RL
Check yourself
0 / 4 answered
- Which verifier is most reliable for 'the refund was issued correctly'?
- What are the parts of an agent environment?
- An agent passes all visible tests by special-casing the example inputs. What catches it?
- Why did Baker et al. warn against training directly against a chain-of-thought monitor?
Spec-Driven Development
Check yourself
0 / 4 answered
- What distinguishes vibe engineering from vibe coding?
- In the Spec Kit workflow, what comes between specify and tasks?
- Why write acceptance criteria as given/when/then cases?
- Where should human review concentrate in spec-driven development, and why?
Q&A Review Bank
Check yourself
0 / 9 answered
- What is an agent harness?
- How do long-running coding harnesses carry work across context windows?
- Why separate generation from evaluation?
- What is a sprint contract?
- When should a harness component be removed?
- When does a recurring task justify building a loop?
- What limits how much a loop can improve an agent's output?
- Name three signals that justify turning a loop into a graph.
- What is the 'dies halfway' test?
Check yourself
0 / 4 answered
- What is context rot?
- Just-in-time vs up-front context?
- Name five techniques for managing a long-running agent's context.
- How should compaction be scheduled given prompt caching?
Check yourself
0 / 5 answered
- What makes a repository agent-ready?
- What did METR's 2025 randomised trial of AI tools find?
- Interactive vs background coding agents?
- API, DOM automation or pixels - how do you choose for a computer-use task?
- Why are browser agents with logged-in sessions high risk?
Check yourself
0 / 5 answered
- What is in a skill, and which fields are required?
- Explain progressive disclosure for skills.
- AGENTS.md, skill, MCP server or memory - where does each kind of knowledge go?
- Why must skills be reviewed like code?
- Name the five policies of a memory subsystem.
Check yourself
0 / 7 answered
- Rank verifier types by reliability.
- What are the parts of an agent environment?
- How did RLVR extend to agents?
- Give three reward-hacking behaviours and a defence for each.
- What did Baker et al. find about chain-of-thought monitoring?
- Vibe coding vs vibe engineering?
- What are the stages of the Spec Kit workflow?
Minimal Coding Harness
Check yourself
0 / 4 answered
- Why are there hidden tests as well as visible ones?
- The stop hook runs the harness's own copy of the visible tests. Which reward hack does that prevent?
- What does the 'claimed-but-failed' column measure?
- Why did adding the stop hook alone not change the results?