Contents
Map

Quiz · 17 · Production Agents

57 questions from 7 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

Production Agent Architecture

Check yourself
0 / 4 answered
  1. An agent's refund tool is retried after a network timeout and the customer is refunded twice. What was missing?
  2. A tool call fails with a validation error on its arguments. What should the harness do?
  3. What should happen when an approval request gets no answer for 24 hours?
  4. Why scale agent workers on queue depth rather than CPU?

Agent Security

Check yourself
0 / 5 answered
  1. An email assistant reads incoming mail, can search the company CRM, and can send email. Which change removes a leg of the lethal trifecta without losing the main use case?
  2. Which pattern keeps untrusted data from changing which tools are called, while still letting it fill in arguments?
  3. In the dual LLM pattern, what does the privileged model receive from the quarantined model?
  4. Why inject credentials at an egress proxy instead of giving the sandbox an API key?
  5. Which OWASP agentic risk covers an attacker planting a false 'fact' that the agent stores and acts on in later sessions?

Durable Execution

Check yourself
0 / 4 answered
  1. Why must a model call be an activity/step rather than code in the workflow body?
  2. A refund step crashed after the payment API succeeded but before the result was recorded. What prevents a double refund on retry?
  3. What does a durable-execution engine add over a LangGraph checkpointer?
  4. An agent's workflow history grows to thousands of large events. What do you do?

Cost and Latency

Check yourself
0 / 4 answered
  1. An agent's input tokens per task grow much faster than its number of turns. Why?
  2. Which change breaks prompt caching for the rest of an agent conversation?
  3. A cheaper model cuts cost per request by 60% but drops task success from 90% to 70%. What happened to cost per successful task?
  4. Name two ways to shrink the per-turn context growth from tool outputs.

Evaluation and Benchmarks

Check yourself
0 / 4 answered
  1. A regression suite passes at 97% after a prompt change, down from 100%. The change improved the capability suite. What do you do?
  2. How do you know an LLM judge is good enough to use?
  3. Why are paired comparisons preferred when comparing two agent versions?
  4. Model A scores higher than model B on SWE-bench Verified. Name three reasons this may not predict which is better for your coding agent.

Agent Observability

Check yourself
0 / 4 answered
  1. Which OpenTelemetry GenAI operation name marks a tool invocation span?
  2. Why is prompt and tool content opt-in in the GenAI conventions?
  3. Approvals for an agent's refunds run at 99.7% with a median review time of 4 seconds. What does the dashboard suggest?
  4. What should you record on every run so a regression can be traced to a change?

Q&A Review Bank

Check yourself
0 / 5 answered
  1. Name the layers of a production agent system.
  2. Which failures should an agent harness retry, and which not?
  3. Why do retries require idempotency keys on writes?
  4. What should happen when a run hits its step or cost bound?
  5. What makes a human approval gate effective rather than ceremonial?
Check yourself
0 / 6 answered
  1. State the lethal trifecta and the structural fix.
  2. Why is prompt-injection resistance in the model not sufficient?
  3. Name the six injection-resistant design patterns.
  4. What does CaMeL add to the dual LLM pattern?
  5. List five risks from the OWASP Top 10 for Agentic Applications.
  6. How should credentials be handled for an agent that runs code?
Check yourself
0 / 5 answered
  1. What does durable execution record, and what does it enable?
  2. What is the determinism rule?
  3. Compare Temporal, Restate and DBOS in one line each.
  4. Is a LangGraph checkpointer durable execution?
  5. How do you keep a long agent's workflow history manageable?
Check yourself
0 / 6 answered
  1. Why does an agent's input-token cost grow faster than its number of turns?
  2. Three rules for prompt caching in an agent loop?
  3. Name four ways to shrink context growth.
  4. Routing vs cascades?
  5. Why track cost per successful task rather than cost per request?
  6. Name four latency levers for agents.
Check yourself
0 / 10 answered
  1. What is trajectory evaluation and why is it needed alongside outcome checks?
  2. How does LLM-as-judge differ for single-turn evaluation and trajectory evaluation?
  3. Regression vs capability suites?
  4. How do you validate an LLM judge?
  5. What do τ-bench's pass^k results show?
  6. Name four questions to ask about a public agent benchmark score.
  7. What are the three core span types in the OpenTelemetry GenAI conventions for agents?
  8. What must you record on every run to tie a regression to a change?
  9. Walk through debugging 'the agent said it refunded but didn't'.
  10. Which signals indicate a possible prompt-injection campaign in production?