Contents
Map

Quiz · 13 · Agent Foundations

37 questions from 9 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

What Are AI Agents

Check yourself
0 / 4 answered
  1. What distinguishes an agent from a workflow in Anthropic's definition?
  2. A support agent drafts refunds and a human must confirm each one before it is issued. Which autonomy level is this?
  3. METR's time-horizon measure reports the task length models complete with what reliability?
  4. A team wants an 'agent' that extracts five fields from invoices and writes them to an ERP. What would you build first, and why?

Anatomy of an AI Agent

Check yourself
0 / 4 answered
  1. An agent keeps calling get_order with ids the user never mentioned. Which component would you fix first?
  2. Why can the same model score very differently on SWE-bench in two different agents?
  3. Which is the most reliable basis for agent self-correction?
  4. Where should a rule like 'never refund more than $500 without approval' be enforced?

The Agent Loop

Check yourself
0 / 5 answered
  1. A Claude response has stop_reason 'max_tokens' and contains a tool_use block. What should the loop do?
  2. What must you do with 'pause_turn'?
  3. A reasoning model returns reasoning items alongside a function_call in the Responses API. What happens to them on the next request?
  4. An episode has 20 model calls and the context grows by about 2,000 tokens per step from a 4,000-token start. Roughly how many input tokens are processed in total?
  5. Why is 'grade on environment state' the fix for claimed actions?

Tool Use & Function Calling

Check yourself
0 / 5 answered
  1. What does strict mode guarantee about tool arguments?
  2. In OpenAI strict mode, how do you express an optional parameter?
  3. Your agent has 180 tools across 12 services and often picks the wrong one. Which change addresses both context cost and selection accuracy?
  4. Why should a payment tool accept an idempotency key?
  5. Rewrite this error for the model: {"error": "E_STATE"}

Agent Memory

Check yourself
0 / 5 answered
  1. A coding agent keeps a list of the steps it has completed in the current task, persisted after each step so it can resume after a crash. Which memory is this?
  2. The agent learns that this team always wants tests written before code and updates its instructions accordingly. Which kind of long-term memory is that?
  3. What are Mem0's four memory-update operations?
  4. Why does Zep store facts with validity intervals instead of overwriting them?
  5. A user asks your assistant to forget everything about them. What must deletion cover?

Planning & Reasoning

Check yourself
0 / 4 answered
  1. According to Huang et al. (2024), when does self-correction of reasoning reliably help?
  2. What is ReWOO's main efficiency gain over ReAct?
  3. Why keep a plan in harness state when the model already reasons internally?
  4. Give two re-planning triggers that are better than 're-plan every step'.

Agents in Practice

Check yourself
0 / 3 answered
  1. Which property most distinguishes coding from many other agent use cases?
  2. An agent's approval gate for sending customer emails has been approved 100% of the time over 5,000 requests. What does that suggest?
  3. What does implementing an 'approve actions' pattern require from the agent loop?

Agent Evaluation Basics

Check yourself
0 / 4 answered
  1. A task succeeded in 2 of 4 trials. What is its pass^2?
  2. Why is pass^k, not pass@k, the right reliability metric for a customer-service agent?
  3. An LLM judge scores 95% of replies as correct, but a database check finds only 60% of tasks completed. What is the most likely explanation?
  4. Why include tasks the agent should refuse in the test set, and how do you grade them?

Agent Loop from Scratch

Check yourself
0 / 3 answered
  1. Why do the cancel_shipped and cancel_delivered tasks pass even in the no_policy variant?
  2. In refund_one, the reply says the refund was processed but the database is unchanged. What class of failure is this, and what grader catches it?
  3. pass^4 is close to pass@1 in every variant. What does that say about this model at temperature 1.0?