Quiz · 13 · Agent Foundations
37 questions from 9 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
What Are AI Agents
Check yourself
0 / 4 answered
- What distinguishes an agent from a workflow in Anthropic's definition?
- A support agent drafts refunds and a human must confirm each one before it is issued. Which autonomy level is this?
- METR's time-horizon measure reports the task length models complete with what reliability?
- A team wants an 'agent' that extracts five fields from invoices and writes them to an ERP. What would you build first, and why?
Anatomy of an AI Agent
Check yourself
0 / 4 answered
- An agent keeps calling get_order with ids the user never mentioned. Which component would you fix first?
- Why can the same model score very differently on SWE-bench in two different agents?
- Which is the most reliable basis for agent self-correction?
- Where should a rule like 'never refund more than $500 without approval' be enforced?
The Agent Loop
Check yourself
0 / 5 answered
- A Claude response has stop_reason 'max_tokens' and contains a tool_use block. What should the loop do?
- What must you do with 'pause_turn'?
- A reasoning model returns reasoning items alongside a function_call in the Responses API. What happens to them on the next request?
- An episode has 20 model calls and the context grows by about 2,000 tokens per step from a 4,000-token start. Roughly how many input tokens are processed in total?
- Why is 'grade on environment state' the fix for claimed actions?
Tool Use & Function Calling
Check yourself
0 / 5 answered
- What does strict mode guarantee about tool arguments?
- In OpenAI strict mode, how do you express an optional parameter?
- Your agent has 180 tools across 12 services and often picks the wrong one. Which change addresses both context cost and selection accuracy?
- Why should a payment tool accept an idempotency key?
- Rewrite this error for the model: {"error": "E_STATE"}
Agent Memory
Check yourself
0 / 5 answered
- A coding agent keeps a list of the steps it has completed in the current task, persisted after each step so it can resume after a crash. Which memory is this?
- The agent learns that this team always wants tests written before code and updates its instructions accordingly. Which kind of long-term memory is that?
- What are Mem0's four memory-update operations?
- Why does Zep store facts with validity intervals instead of overwriting them?
- A user asks your assistant to forget everything about them. What must deletion cover?
Planning & Reasoning
Check yourself
0 / 4 answered
- According to Huang et al. (2024), when does self-correction of reasoning reliably help?
- What is ReWOO's main efficiency gain over ReAct?
- Why keep a plan in harness state when the model already reasons internally?
- Give two re-planning triggers that are better than 're-plan every step'.
Agents in Practice
Check yourself
0 / 3 answered
- Which property most distinguishes coding from many other agent use cases?
- An agent's approval gate for sending customer emails has been approved 100% of the time over 5,000 requests. What does that suggest?
- What does implementing an 'approve actions' pattern require from the agent loop?
Agent Evaluation Basics
Check yourself
0 / 4 answered
- A task succeeded in 2 of 4 trials. What is its pass^2?
- Why is pass^k, not pass@k, the right reliability metric for a customer-service agent?
- An LLM judge scores 95% of replies as correct, but a database check finds only 60% of tasks completed. What is the most likely explanation?
- Why include tasks the agent should refuse in the test set, and how do you grade them?
Agent Loop from Scratch
Check yourself
0 / 3 answered
- Why do the cancel_shipped and cancel_delivered tasks pass even in the no_policy variant?
- In refund_one, the reply says the refund was processed but the database is unchanged. What class of failure is this, and what grader catches it?
- pass^4 is close to pass@1 in every variant. What does that say about this model at temperature 1.0?