These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Check yourself 0 / 4 answered
An open-weight instruct model gives rambling, off-format answers when you call it through a raw text-completion endpoint, but works well in the provider's chat playground. What is the most likely cause?
A The prompt was not rendered with the model's chat template B Temperature was too high C The context window was exceeded D The model needs few-shot examples Why can a user message override instructions in a system prompt?
A System-prompt priority is a trained tendency - all roles become tokens read by the same attention layers B The system prompt is truncated when the user message is long C APIs send the user message before the system prompt D System prompts are only applied on the first turn You need byte-identical outputs for a regression test. Which statement is correct?
A Setting temperature to 0 guarantees identical outputs B Even at temperature 0, batching and floating-point reduction order can change outputs; pin the model and compare on eval sets, or use batch-invariant inference you control C Setting top_p to 1 guarantees identical outputs D Only reasoning models are non-deterministic Your code prefills the assistant turn with '{' to force JSON, and after a model upgrade the API returns a 400 error. What is the recommended replacement?
Show answer
Check yourself 0 / 4 answered
You give a large frontier model 8 few-shot sentiment examples in which every label is deliberately flipped (positive reviews labelled negative). What does current evidence predict?
A The model largely follows the flipped labels, overriding its prior knowledge B The model ignores the labels and answers correctly anyway C The model refuses the task D Accuracy is unchanged because only the format matters For which task is adding 'Let's think step by step' to a non-reasoning model most likely to help?
A A multi-step word problem requiring arithmetic B Retrieving the capital of a country C Classifying the sentiment of a one-line review D Translating a sentence In a ReAct loop, who executes the tool?
A The application code - the model only emits a tool call (as text or as a structured block) B The model, inside its forward pass C The provider, always D The tokenizer Your few-shot prompt scored 3 points higher than zero-shot on 50 examples. What should you do before shipping it?
Show answer
Check yourself 0 / 4 answered
Which task is self-consistency with majority voting best suited to?
A Multi-step math word problems with a single numeric answer B Writing a product description C Summarising a meeting D Brainstorming slogans A self-refine loop on math problems (no tools, no tests) lowers accuracy. What is the most likely explanation?
A Without external feedback, models change correct answers as often as wrong ones B The temperature was too low C The context window overflowed D Self-refine only works for reasoning models What does Tree of Thoughts need in order to beat chain-of-thought reliably?
A A trustworthy way to evaluate partial solutions B A larger context window C Few-shot examples of trees D A lower temperature A 3-step prompt chain produces bad summaries. How do you find which step is at fault?
Show answer
Check yourself 0 / 4 answered
JSON mode is enabled, yet your parser rejects some responses with a KeyError. Why?
A JSON mode guarantees valid JSON syntax, not your schema's keys and types B JSON mode only works at temperature 0 C The tokenizer dropped the key D JSON mode requires few-shot examples How does a constrained decoder prevent invalid output?
A It masks the logits of tokens that cannot legally follow in the grammar's current state, so they can't be sampled B It retries until the output parses C It fine-tunes the model on the schema D It post-processes the text with regexes A schema is {"answer": string, "explanation": string} and accuracy on a reasoning task dropped after adding it to a non-reasoning model. What is the cheapest fix to try first?
A Reorder the fields so a reasoning/explanation field comes before the answer B Remove the schema and parse free text C Increase the temperature D Add more few-shot examples Your structured-output call returns text that fails JSON parsing even though constrained decoding is on. Name two likely causes.
Show answer
Check yourself 0 / 4 answered
You migrate a classification prompt from a non-reasoning model to a reasoning model. Which change is most likely to help?
A Delete the 'think step by step' instructions and procedure, keep the label definitions and output schema, and try a low effort setting B Add more worked reasoning examples C Raise the temperature D Add 'CRITICAL' before every rule A reasoning model's answers are cut off mid-JSON although the answer itself is short. What is the likely cause?
A max_tokens covers reasoning tokens too, and the budget was exhausted by thinking B The schema is too long C Reasoning models cannot produce JSON D The input was too long In a tool-calling loop you strip the thinking blocks from the assistant message before sending the tool results back. What happens?
A Quality degrades or the API rejects the request, because the reasoning state needed to continue is lost B Nothing - thinking blocks are only for display C The model re-reads its reasoning from its cache automatically D Latency improves with no side effects Why shouldn't you treat a model's reasoning summary as the explanation for its decision in an audit?
Show answer
Check yourself 0 / 4 answered
An assistant's answers get worse over a long session even though the context never exceeds the model's window. Which strategy most directly addresses this?
A Compress: compact older turns and clear stale tool results B Switch to a model with a larger window C Raise the temperature D Move the system prompt to the end What did the Lost in the Middle study find for GPT-3.5-Turbo with 20-30 retrieved documents and the answer in the middle?
A Accuracy fell below its closed-book accuracy (no documents at all) B Accuracy was unaffected by position C Accuracy was highest in the middle D The model refused to answer Which is an example of the 'isolate' strategy?
A A sub-agent explores a codebase in its own context and returns a short summary to the main agent B Summarising the first 50 turns of a conversation C Retrieving the top 5 documents D Saving user preferences to a database For a long contract plus a question, where should the question go and why?
Show answer
Check yourself 0 / 4 answered
You add the current date and time to the first line of your system prompt. What happens to prompt caching?
A Every request misses the cache, because the prefix changes on each call B Nothing - caching ignores the system prompt C Only the date line is recomputed D The cache hit rate doubles With a 1.25× write price and 0.1× read price, how many requests sharing a prefix within the TTL are needed for caching to be cheaper than no caching?
A 1 B 2 C 5 D 10 What does a prompt cache store?
A The KV tensors computed for a prompt prefix, so prefill for that prefix can be skipped B The model's previous responses C Embeddings of the prompt D A compressed copy of the model weights An agent's cache hit rate fell from 85% to 10% after a release. Name two likely causes to check first.
Show answer
Check yourself 0 / 4 answered
Which attack does not require the attacker to interact with your application at all?
A Indirect prompt injection via a web page your agent reads B A jailbreak typed into the chat box C Prompt leaking via 'repeat everything above' D A denial-of-service flood of requests What is the most effective way to limit the damage of a successful prompt injection in an email assistant?
A Least privilege and human confirmation before sending or forwarding email B A longer system prompt forbidding injections C A regex filter for 'ignore previous instructions' D Using temperature 0 A new prompt scores 91.5% vs 90.0% for the old one on 200 items; the paired 95% CI of the difference is [-1.0%, +4.0%]. What should you conclude?
A The improvement is not established; the interval includes zero B Ship it - it is 1.5 points better C The new prompt is worse D The eval set is too large Why pin dated model snapshots rather than an alias like 'latest' in production?
Show answer
Check yourself 0 / 4 answered
What is the main difference between OPRO and GEPA?
A OPRO learns from a history of prompt scores; GEPA also reads execution traces and textual feedback and reflects on failures B OPRO fine-tunes weights; GEPA does not C GEPA only works with few-shot examples D OPRO requires no metric Your DSPy-optimized prompt scores 94% on the validation set used during optimization and 81% on new production data. What is the most likely cause?
A Overfitting to a small or unrepresentative optimization set B The model's temperature changed C DSPy does not support classification D The prompt was too short In DSPy, what does the optimizer need from you that a hand-written prompt doesn't?
A A metric and labelled (or judge-scorable) examples B A GPU C The exact prompt text D A vector database When would BootstrapFinetune be a better choice than instruction optimization?
Show answer
Check yourself 0 / 3 answered
Variant A scores 78% and variant B 75% on the same 52 items; the paired 95% CI of A−B is [−5%, +12%]. What do you report?
A No established difference - collect more data or accept either B A is 3 points better C B is better because its interval is narrower D The experiment is invalid Why did schema-constrained decoding not improve the valid-output rate in this lab?
A The schema is a single enum field, which the instruct model already produced reliably B Outlines was not actually applied C The model was too large D Temperature was too high Why does the script shuffle the test set with a fixed seed before applying --limit?
Show answer