Contents
Map

11 ยท Prompt & Context Engineering

Automated Prompt Optimization

View as:

Automated Prompt Optimization

Automated prompt optimization treats prompt text and few-shot examples as parameters to search over, scoring candidates with a metric on a training set - replacing manual trial and error with a reproducible optimization loop.

Learning objectives 55 min
By the end of this page you will be able to:
  • Describe the generate-evaluate-select loop shared by APE, OPRO, MIPROv2 and GEPA, and what distinguishes each
  • Write a DSPy program with a signature, a module and a metric, and compile it with an optimizer
  • Choose between few-shot bootstrapping, instruction search, reflective evolution and fine-tuning for a pipeline
  • Avoid overfitting the optimizer to a small training set
Prerequisites

The Loop

Every method in this note is a variation on the same loop. What differs is how candidates are proposed.

flowchart LR
    P["๐Ÿงช Propose candidates<br/>instructions, examples"] --> E["๐Ÿ“ Evaluate<br/>metric on a training batch"]
    E --> S["๐Ÿ† Select / update<br/>keep the best, learn from scores<br/>or from failure traces"]
    S --> P
    S --> V["โœ… Validate on held-out data<br/>before shipping"]

    style P fill:#e8e0d4,stroke:#c8b89a
    style E fill:#d8dfe8,stroke:#b0bac8
    style S fill:#dde4dc,stroke:#b0c4b0
    style V fill:#ddd8e4,stroke:#b8b0c8
MethodHow it proposesSignal it learns from
APE (Zhou et al., 2023)An LLM infers candidate instructions from input-output examplesScores; pick the best (one round, optionally resampled around the winner)
OPRO (Yang et al., 2024)An optimizer LLM reads a history of (prompt, score) pairs and proposes a better promptThe score trajectory
DSPy bootstrappingRuns the program on training inputs and keeps traces that pass the metric as few-shot demosWhich demos lead to passing outputs
MIPROv2 (Opsahl-Ong et al., 2024)Proposes instructions grounded in the program, the data and bootstrapped demos, then searches instruction ร— demo combinations with Bayesian optimizationScores on minibatches
GEPA (Agrawal et al., 2025)Reads full execution traces and textual feedback, reflects on what went wrong, and evolves instructions; keeps a Pareto front of candidates that win on different examplesScores and natural-language feedback
TextGrad (Yuksekgonul et al., 2025)Treats LLM-written critiques as "textual gradients" and back-propagates them through a computation graphCritiques of outputs

GEPA's authors report that reflective prompt evolution can match or beat reinforcement-learning fine-tuning (GRPO) on several tasks with far fewer rollouts, and outperforms MIPROv2 - evidence that for many pipelines, optimizing the prompt is a cheaper first move than training weights.


DSPy: Programs, Not Prompts

DSPy (Khattab et al., 2024) separates what each LM call should do from how it is prompted. You declare signatures (typed inputs and outputs), compose modules into a program, and let an optimizer write the actual prompts and choose examples against your metric.

import dspy

dspy.configure(lm=dspy.LM("openai/<model>"))  # any LiteLLM-style "provider/model" string

class TicketClassifier(dspy.Signature):
    """Classify a customer-support ticket."""
    ticket: str = dspy.InputField()
    category: str = dspy.OutputField(desc="billing | bug | account_access | feature_request")

classify = dspy.ChainOfThought(TicketClassifier)   # adds a reasoning field before the output
print(classify(ticket="I was charged twice this month.").category)

A multi-step program is ordinary Python; retrieval is just a function you call:

class SupportAnswer(dspy.Module):
    def __init__(self, search):
        self.search = search                                         # any retriever function
        self.answer = dspy.ChainOfThought("context, question -> answer")

    def forward(self, question):
        context = self.search(question, k=5)
        return self.answer(context=context, question=question)

Optimizing it

trainset = [dspy.Example(ticket=t, category=c).with_inputs("ticket") for t, c in train_pairs]
valset   = [dspy.Example(ticket=t, category=c).with_inputs("ticket") for t, c in val_pairs]

def metric(gold, pred, trace=None):
    return gold.category == pred.category

optimizer = dspy.MIPROv2(metric=metric, auto="light")       # "light" | "medium" | "heavy" budget
optimized = optimizer.compile(classify, trainset=trainset, valset=valset)
optimized.save("ticket_classifier.json")                     # instructions + demos, versionable

# Reflective evolution: the metric can return feedback text as well as a score
def metric_with_feedback(gold, pred, trace=None, pred_name=None, pred_trace=None):
    ok = gold.category == pred.category
    feedback = "correct" if ok else f"expected {gold.category}, got {pred.category}"
    return dspy.Prediction(score=float(ok), feedback=feedback)

gepa = dspy.GEPA(metric=metric_with_feedback, auto="light",
                 reflection_lm=dspy.LM("openai/<strong-model>"))
evolved = gepa.compile(classify, trainset=trainset, valset=valset)
DSPy optimizerWhat it changesWhen to use
BootstrapFewShot / BootstrapFewShotWithRandomSearchFew-shot demosSmall data (tens of examples), quick wins
MIPROv2Instructions and demos jointlyHundreds of examples, multi-module programs
GEPAInstructions, driven by traces and textual feedbackWhen you can explain failures in words; sample-efficient
BootstrapFinetuneThe model weights, from the program's own successful tracesHigh volume, stable task, a smaller model you can fine-tune

When DSPy fits

SituationFit
Multi-step pipeline (retrieve, reason, extract) with a clear metricStrong
You need to move a pipeline to a cheaper model and re-tune itStrong - recompile against the new model
A single simple prompt, no labelled dataWeak - write it by hand
No programmatic or judge-based metric existsWeak - write the metric first; it is usually the hard part
You need exact, human-readable control of every promptMixed - compiled prompts can be inspected, saved and reviewed, but they are generated

Pitfalls

  • The metric is the specification. An optimizer maximises exactly what you measure. A metric that checks only format will produce well-formatted wrong answers; a lenient LLM judge will be gamed. Validate the metric on hand-labelled examples first.
  • Overfitting. With a few dozen training examples, optimizers find prompts that fit those examples. Keep a held-out test set the optimizer never sees and report results on it - never on the training or validation set used during search.
  • Model coupling. An optimized prompt is tuned to one model. Re-run optimization (not just evaluation) when you change models.
  • Cost. Optimization runs hundreds or thousands of LM calls. Use the budget presets, cache LM calls, and use a cheaper model for trial runs.

Check Yourself

Check yourself
0 / 4 answered
  1. What is the main difference between OPRO and GEPA?
  2. Your DSPy-optimized prompt scores 94% on the validation set used during optimization and 81% on new production data. What is the most likely cause?
  3. In DSPy, what does the optimizer need from you that a hand-written prompt doesn't?
  4. When would BootstrapFinetune be a better choice than instruction optimization?

Exercises

Exercise - Optimize the lab classifier

Port the code lab's ticket classifier to DSPy. Split the 60 tickets into 20 train / 12 validation / 28 test. Compare zero-shot dspy.Predict, BootstrapFewShot and MIPROv2(auto="light") on the test set with bootstrap confidence intervals. Use a local model through an OpenAI-compatible server (e.g. vLLM or Ollama) or any API model.

Hint

dspy.LM accepts api_base= for a local OpenAI-compatible endpoint.

Solution

Expect bootstrapped demos to help a small model noticeably and MIPROv2 to add little beyond that on such a small, simple task - with wide intervals on 28 test items. The honest conclusion is usually "not distinguishable at this sample size"; the value of the exercise is the workflow and the held-out discipline.

Exercise - Write a metric that can't be gamed

For a RAG answer-generation step, write a DSPy metric that rewards answers that are correct and supported by the retrieved context, and returns textual feedback for GEPA. Describe how you would check the metric itself before optimizing.

Solution

Combine an exact/semantic match against the reference answer with a groundedness check (an LLM judge asked whether each claim is supported by the context), returning feedback such as "claim X not supported by context". Before optimizing, score 50 hand-labelled outputs with the metric and compare to human judgements; fix disagreements (e.g. the judge accepting unsupported claims) before letting an optimizer exploit them.

Study Notes

Must-know:

  • All methods: propose โ†’ evaluate on a metric โ†’ select; validate on held-out data
  • APE (infer instructions), OPRO (score history), MIPROv2 (instruction + demo search with Bayesian optimization), GEPA (reflective evolution from traces and feedback), TextGrad (textual gradients)
  • DSPy: signatures + modules + metric + optimizer; compiled programs can be saved and versioned
  • The metric is the spec; overfitting and model coupling are the main risks
  • Prompt optimization is often a cheaper first step than RL fine-tuning

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท