Quiz · 17 · Production Agents
57 questions from 7 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Production Agent Architecture
Check yourself
0 / 4 answered
- An agent's refund tool is retried after a network timeout and the customer is refunded twice. What was missing?
- A tool call fails with a validation error on its arguments. What should the harness do?
- What should happen when an approval request gets no answer for 24 hours?
- Why scale agent workers on queue depth rather than CPU?
Agent Security
Check yourself
0 / 5 answered
- An email assistant reads incoming mail, can search the company CRM, and can send email. Which change removes a leg of the lethal trifecta without losing the main use case?
- Which pattern keeps untrusted data from changing which tools are called, while still letting it fill in arguments?
- In the dual LLM pattern, what does the privileged model receive from the quarantined model?
- Why inject credentials at an egress proxy instead of giving the sandbox an API key?
- Which OWASP agentic risk covers an attacker planting a false 'fact' that the agent stores and acts on in later sessions?
Durable Execution
Check yourself
0 / 4 answered
- Why must a model call be an activity/step rather than code in the workflow body?
- A refund step crashed after the payment API succeeded but before the result was recorded. What prevents a double refund on retry?
- What does a durable-execution engine add over a LangGraph checkpointer?
- An agent's workflow history grows to thousands of large events. What do you do?
Cost and Latency
Check yourself
0 / 4 answered
- An agent's input tokens per task grow much faster than its number of turns. Why?
- Which change breaks prompt caching for the rest of an agent conversation?
- A cheaper model cuts cost per request by 60% but drops task success from 90% to 70%. What happened to cost per successful task?
- Name two ways to shrink the per-turn context growth from tool outputs.
Evaluation and Benchmarks
Check yourself
0 / 4 answered
- A regression suite passes at 97% after a prompt change, down from 100%. The change improved the capability suite. What do you do?
- How do you know an LLM judge is good enough to use?
- Why are paired comparisons preferred when comparing two agent versions?
- Model A scores higher than model B on SWE-bench Verified. Name three reasons this may not predict which is better for your coding agent.
Agent Observability
Check yourself
0 / 4 answered
- Which OpenTelemetry GenAI operation name marks a tool invocation span?
- Why is prompt and tool content opt-in in the GenAI conventions?
- Approvals for an agent's refunds run at 99.7% with a median review time of 4 seconds. What does the dashboard suggest?
- What should you record on every run so a regression can be traced to a change?
Q&A Review Bank
Check yourself
0 / 5 answered
- Name the layers of a production agent system.
- Which failures should an agent harness retry, and which not?
- Why do retries require idempotency keys on writes?
- What should happen when a run hits its step or cost bound?
- What makes a human approval gate effective rather than ceremonial?
Check yourself
0 / 6 answered
- State the lethal trifecta and the structural fix.
- Why is prompt-injection resistance in the model not sufficient?
- Name the six injection-resistant design patterns.
- What does CaMeL add to the dual LLM pattern?
- List five risks from the OWASP Top 10 for Agentic Applications.
- How should credentials be handled for an agent that runs code?
Check yourself
0 / 5 answered
- What does durable execution record, and what does it enable?
- What is the determinism rule?
- Compare Temporal, Restate and DBOS in one line each.
- Is a LangGraph checkpointer durable execution?
- How do you keep a long agent's workflow history manageable?
Check yourself
0 / 6 answered
- Why does an agent's input-token cost grow faster than its number of turns?
- Three rules for prompt caching in an agent loop?
- Name four ways to shrink context growth.
- Routing vs cascades?
- Why track cost per successful task rather than cost per request?
- Name four latency levers for agents.
Check yourself
0 / 10 answered
- What is trajectory evaluation and why is it needed alongside outcome checks?
- How does LLM-as-judge differ for single-turn evaluation and trajectory evaluation?
- Regression vs capability suites?
- How do you validate an LLM judge?
- What do τ-bench's pass^k results show?
- Name four questions to ask about a public agent benchmark score.
- What are the three core span types in the OpenTelemetry GenAI conventions for agents?
- What must you record on every run to tie a regression to a change?
- Walk through debugging 'the agent said it refunded but didn't'.
- Which signals indicate a possible prompt-injection campaign in production?