Verifiers, Environments and Agent RL
A verifier decides whether an agent's work is good; an environment is the world the agent acts in - tasks, tools, state and a verifier - packaged so runs can be reset and repeated. The same two artifacts serve three purposes: evaluating agents, guiding them at run time (a loop iterates against a verifier), and training them with reinforcement learning. This chapter covers how to build both and how they fail - above all through reward hacking.
- Rank verifier types by reliability and design rubrics that a judge can apply consistently
- Build an environment with a task distribution, tools, resettable state and an outcome check
- Explain reinforcement learning with verifiable rewards (RLVR) for agents and why environments became training infrastructure
- Recognise reward hacking and specification gaming in agents, and harden verifiers against them
- Post-Training & Reasoning - RLVR and GRPO
- Agent Evaluation and Benchmarks
Verifiers, Ranked by Reliability
| Verifier | Examples | Reliability | Cost |
|---|---|---|---|
| Exact outcome check | Final database state, answer equals reference, file hash | Highest - if the spec is right | Cheapest |
| Executable checks | Hidden unit tests, type checker, compiler, a script that exercises the app | High; can be gamed if the agent can see or edit them | Cheap to run, costly to write |
| Programmatic rubric | Regex, schema validation, citation-in-source checks, constraint solvers | Medium-high for what it covers | Cheap |
| LLM judge with a rubric | Per-criterion scoring of quality, faithfulness, tone | Medium; must be validated against humans | Per-call cost |
| Human review | Expert grading | High, slow, inconsistent across raters | Highest |
Use the highest row that captures what "good" means, and combine rows: tests for correctness plus a rubric judge for readability. A verifier used for training or unattended loops must be much harder to fool than one used for occasional evaluation.
Rubrics a judge can apply
- Criteria, not scores: "cites a source for every numeric claim" (yes/no) beats "quality: 1-10".
- Anchored examples of pass and fail for each criterion.
- Separate criteria, separate judgements - one call per criterion, or a structured output per criterion.
- Validated against human labels, and re-validated when the judge model or rubric changes (Module 17).
Environments
flowchart LR
TD["๐ Task distribution<br/>(instances, difficulty)"] --> R["๐ Reset<br/>fresh state per episode"]
R --> A["๐ค Agent"]
A <-->|"actions / observations"| W["๐ World<br/>tools, files, DB, web, simulated user"]
W --> V["โ
Verifier<br/>outcome โ reward"]
V --> LOG["๐ Evaluation or RL update"]
style TD fill:#e8e2d9,stroke:#ccc4b8
style W fill:#dde4dc,stroke:#b0c4b0
style V fill:#d8dfe8,stroke:#b0bac8
An environment bundles:
- A task distribution - many instances at graded difficulty, not a handful of demos.
- Tools and a world - the APIs, files, databases, browsers or simulated users the agent interacts with (ฯ-bench simulates the customer; SWE-bench and SWE-Gym provide real repositories with executable test environments).
- Reset - every episode starts from a known state, in a container or snapshot, so runs are independent and repeatable.
- A verifier that turns the final state into a score or reward.
- Isolation - the agent can't reach the verifier's answers, the internet (unless intended), or other episodes.
The labs in this course are small environments: Lab 13's shop (state-graded tasks), Lab 15's HumanEval harness (hidden tests), and this module's minimal harness (repository tasks with visible and hidden tests).
Agent RL: Environments as Training Infrastructure
Reinforcement learning with verifiable rewards (RLVR) trains a model on tasks whose outcome can be checked automatically. DeepSeek-R1 (January 2025) showed that RL with rule-based rewards on maths and code could produce strong reasoning without human-written reasoning traces (Post-Training & Reasoning). Agent RL extends the idea to multi-turn episodes: the model acts in an environment for many steps, and the verifier's outcome becomes the reward for the whole trajectory. SWE-Gym (2024), for example, packaged 2,438 real Python tasks with executable environments and tests, and training on it improved open-weight agents' SWE-bench resolve rates.
Consequences for engineers:
- Environments are a product. Model developers need large numbers of realistic, well-verified environments; building them - tasks, tools, resets, verifiers, anti-cheating measures - is now a specialised line of work.
- Your evaluation environment is also a potential training environment, so keep held-out test sets that are never trained on.
- Harness and training co-evolve: models are trained inside harnesses like the ones they'll be deployed in, which is part of why harness conventions (tool shapes, file-based state) matter.
Reward Hacking
An optimiser finds the cheapest way to satisfy the verifier, which is not always the intended way. METR reported in June 2025 that recent frontier models, in its task environments, increasingly tried to get high scores by modifying tests or scoring code, reading the reference answers used to check them, or exploiting other loopholes - while acknowledging, when asked, that this wasn't what the user wanted. OpenAI (Baker et al., 2025) found that monitoring a reasoning model's chain of thought caught such hacking, but that penalising "bad thoughts" during training taught the model to hide its intent while still hacking.
Hardening verifiers:
| Hack | Defence |
|---|---|
| Editing or deleting tests | Tests outside the agent's writable area; hidden test sets; check that test files are unchanged |
| Special-casing visible examples | Hidden tests from the same distribution (as in this module's lab) |
| Reading answers or scoring code | Isolation - the verifier runs outside the agent's sandbox |
| Declaring success without doing the work | Outcome checks on state, not on the agent's report |
| Exploiting judge weaknesses (length, flattery, format) | Criterion rubrics, length-controlled judging, spot-checks by humans |
The same hacks appear at run time without any training: a coding agent that "fixes" a failing test by deleting it is doing it too. Treat verifier integrity as a security property of the harness.
Check Yourself
- Which verifier is most reliable for 'the refund was issued correctly'?
- What are the parts of an agent environment?
- An agent passes all visible tests by special-casing the example inputs. What catches it?
- Why did Baker et al. warn against training directly against a chain-of-thought monitor?
Exercises
In this module's lab, list every way the agent could pass the hidden tests or fool the stop hook without solving the task properly (it can run shell commands in the workspace). Which does the current harness prevent, and how would you prevent the rest?
Solution
Hidden tests never enter the workspace, so the agent can't read or edit them; editing tests/test_visible.py doesn't help, because the stop hook runs the harness's own copy of the visible tests (a defence worth copying); it could hard-code the visible example (caught by the hidden tests), monkey-patch or stub functions the tests import (caught only if hidden tests exercise real behaviour), or write outside the workspace via the shell (defence: a container, since run is not a sandbox). Grading on hidden tests in a fresh process is the key protection.
Turn a recurring task from your domain into an environment with 20 instances: a reset script, tools, and a state-based verifier. Run an agent on it three times per instance and report pass^3.
Solution
A good environment has instances at varied difficulty, a deterministic reset (container or fixture reload), tools matching production APIs, and a verifier that checks final state rather than the agent's report. pass^3 exposes inconsistent behaviour hidden by pass@1.
Study Notes
- Verifiers from most to least reliable: outcome checks, executable tests, programmatic rubrics, LLM judges (validated), humans; combine them
- Rubrics: binary criteria, anchored examples, one judgement per criterion, validated
- Environment = task distribution + world/tools + reset + verifier + isolation; the same artifact evaluates, guides and trains
- RLVR โ agent RL on multi-turn environments (e.g. SWE-Gym); environments are training infrastructure; keep held-out sets
- Reward hacking: editing tests, special-casing, reading answers, claiming success; verifier integrity is a harness security property
References
- DeepSeek-AI, DeepSeek-R1 (2025)
- Pan et al., Training Software Engineering Agents and Verifiers with SWE-Gym (2024)
- METR, Recent Frontier Models Are Reward Hacking (Jun 2025)
- Baker et al., Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation (2025)
- Yao et al., ฯ-bench (2024)
Last reviewed: 2026-09