Contents
Map

Quiz · 11 · Prompt & Context Engineering

39 questions from 10 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

Prompt Fundamentals

Check yourself
0 / 4 answered
  1. An open-weight instruct model gives rambling, off-format answers when you call it through a raw text-completion endpoint, but works well in the provider's chat playground. What is the most likely cause?
  2. Why can a user message override instructions in a system prompt?
  3. You need byte-identical outputs for a regression test. Which statement is correct?
  4. Your code prefills the assistant turn with '{' to force JSON, and after a model upgrade the API returns a 400 error. What is the recommended replacement?

Core Techniques

Check yourself
0 / 4 answered
  1. You give a large frontier model 8 few-shot sentiment examples in which every label is deliberately flipped (positive reviews labelled negative). What does current evidence predict?
  2. For which task is adding 'Let's think step by step' to a non-reasoning model most likely to help?
  3. In a ReAct loop, who executes the tool?
  4. Your few-shot prompt scored 3 points higher than zero-shot on 50 examples. What should you do before shipping it?

Advanced Techniques

Check yourself
0 / 4 answered
  1. Which task is self-consistency with majority voting best suited to?
  2. A self-refine loop on math problems (no tools, no tests) lowers accuracy. What is the most likely explanation?
  3. What does Tree of Thoughts need in order to beat chain-of-thought reliably?
  4. A 3-step prompt chain produces bad summaries. How do you find which step is at fault?

Structured Outputs

Check yourself
0 / 4 answered
  1. JSON mode is enabled, yet your parser rejects some responses with a KeyError. Why?
  2. How does a constrained decoder prevent invalid output?
  3. A schema is {"answer": string, "explanation": string} and accuracy on a reasoning task dropped after adding it to a non-reasoning model. What is the cheapest fix to try first?
  4. Your structured-output call returns text that fails JSON parsing even though constrained decoding is on. Name two likely causes.

Prompting Reasoning Models

Check yourself
0 / 4 answered
  1. You migrate a classification prompt from a non-reasoning model to a reasoning model. Which change is most likely to help?
  2. A reasoning model's answers are cut off mid-JSON although the answer itself is short. What is the likely cause?
  3. In a tool-calling loop you strip the thinking blocks from the assistant message before sending the tool results back. What happens?
  4. Why shouldn't you treat a model's reasoning summary as the explanation for its decision in an audit?

Context Engineering

Check yourself
0 / 4 answered
  1. An assistant's answers get worse over a long session even though the context never exceeds the model's window. Which strategy most directly addresses this?
  2. What did the Lost in the Middle study find for GPT-3.5-Turbo with 20-30 retrieved documents and the answer in the middle?
  3. Which is an example of the 'isolate' strategy?
  4. For a long contract plus a question, where should the question go and why?

Prompt Caching & Cost

Check yourself
0 / 4 answered
  1. You add the current date and time to the first line of your system prompt. What happens to prompt caching?
  2. With a 1.25× write price and 0.1× read price, how many requests sharing a prefix within the TTL are needed for caching to be cheaper than no caching?
  3. What does a prompt cache store?
  4. An agent's cache hit rate fell from 85% to 10% after a release. Name two likely causes to check first.

Prompts in Production

Check yourself
0 / 4 answered
  1. Which attack does not require the attacker to interact with your application at all?
  2. What is the most effective way to limit the damage of a successful prompt injection in an email assistant?
  3. A new prompt scores 91.5% vs 90.0% for the old one on 200 items; the paired 95% CI of the difference is [-1.0%, +4.0%]. What should you conclude?
  4. Why pin dated model snapshots rather than an alias like 'latest' in production?

Automated Prompt Optimization

Check yourself
0 / 4 answered
  1. What is the main difference between OPRO and GEPA?
  2. Your DSPy-optimized prompt scores 94% on the validation set used during optimization and 81% on new production data. What is the most likely cause?
  3. In DSPy, what does the optimizer need from you that a hand-written prompt doesn't?
  4. When would BootstrapFinetune be a better choice than instruction optimization?

Prompt Evals & Structured Outputs

Check yourself
0 / 3 answered
  1. Variant A scores 78% and variant B 75% on the same 52 items; the paired 95% CI of A−B is [−5%, +12%]. What do you report?
  2. Why did schema-constrained decoding not improve the valid-output rate in this lab?
  3. Why does the script shuffle the test set with a fixed seed before applying --limit?