Contents
Map

18 ยท Agent Engineering

Coding Agents

View as:

Coding Agents

Coding agents - Claude Code, OpenAI Codex, Gemini CLI, Cursor's and GitHub Copilot's agent modes, and many open-source equivalents - are the most widely used agents today and the clearest case study in harness engineering. They are a model in a loop with file, search and shell tools over a repository, verifying their work with the project's own tests. This chapter dissects how they work, how to set a repository up for them, and what the evidence says about their productivity.

Learning objectives 45 min
By the end of this page you will be able to:
  • Describe the architecture shared by coding agents - tools, repository context, permissions and sandbox, verification, git
  • Set up a repository for agents with AGENTS.md, fast deterministic checks and clear task specifications
  • Choose between local interactive agents, background (cloud) agents and parallel agents in worktrees for a task
  • Interpret coding-agent benchmarks and productivity studies, including the METR randomised trial

Anatomy

flowchart LR
    U["๐Ÿ‘ค Task<br/>(issue, prompt, spec)"] --> L["๐Ÿ” Agent loop"]
    L --> T["๐Ÿ› ๏ธ Tools<br/>read, search, edit, write,<br/>shell, web fetch, MCP"]
    T --> R["๐Ÿ“ Repository<br/>in a sandbox or worktree"]
    R -->|"test, lint, type-check output"| L
    L --> G["๐ŸŒฑ git<br/>diff, commit, PR"]
    C["๐Ÿ“œ AGENTS.md / CLAUDE.md<br/>skills, settings"] --> L
    P["๐Ÿ” Permissions<br/>allow / ask / deny,<br/>sandbox, network"] -.-> T

    style L fill:#d8dfe8,stroke:#b0bac8
    style T fill:#e8e0d4,stroke:#c8b89a
    style R fill:#dde4dc,stroke:#b0c4b0
    style P fill:#ddd8e4,stroke:#b8b0c8
ComponentIn practice
ToolsFile read/search (grep, glob), exact string edits rather than whole-file rewrites, file writes, a shell, web fetch, MCP servers for issue trackers, docs and databases
Repository contextLoaded just in time by search and reads, plus an instructions file read at start - AGENTS.md (an open format released by OpenAI in August 2025, now an Agentic AI Foundation project), or CLAUDE.md for Claude Code
Permissions and sandboxRules for which commands run without asking; OS-level sandboxes that restrict file writes to the workspace and block or allowlist network access
VerificationThe project's tests, linters and type checkers - run by the agent, and ideally enforced by the harness before it reports done
GitEvery change is a reviewable diff; commits are checkpoints; worktrees give parallel agents isolated copies
Context managementCompaction for long sessions, sub-agents for exploration, to-do lists for multi-step plans

The loop is the one from Module 13; the differences between products are in the harness - tool design, permission model, context management, and how they verify.

Modes of Use

ModeExamplesGood forWatch out for
Interactive, localClaude Code, Codex CLI, Gemini CLI in a terminal; IDE agent modesExploratory work, debugging, changes you want to steerYour attention is the bottleneck; approve commands carefully
Background / cloudCodex cloud tasks, GitHub Copilot coding agent, Claude Code on the webWell-specified issues; many small tasks in parallel; results as PRsSpecification quality decides outcome; review burden moves to PRs
Parallel local agentsSeveral sessions in separate git worktreesIndependent tasks at once; trying alternative approachesMerge conflicts; review capacity
Headless / CIAgent SDKs and CLIs in pipelines (claude -p, codex exec)Automated fixes, triage, code review on every PRScoped credentials; never give CI agents production secrets

Setting a Repository Up for Agents

The OpenAI and Anthropic harness write-ups agree on the essentials: agents do well in repositories that make the right thing easy to find and easy to check.

  1. An AGENTS.md with the commands to build, test and lint; conventions that aren't obvious from the code; where things live; what not to touch. Keep it short and current; link to deeper docs.
  2. Fast, deterministic checks - unit tests that run in seconds, a type checker, a linter, a formatter. These are the agent's feedback signal; slow or flaky tests make every loop worse.
  3. Clear task specifications - acceptance criteria, examples, the files involved (Spec-Driven Development).
  4. Mechanical enforcement of architecture - lint rules and structural tests for layering and naming, so the agent gets an error instead of a code-review comment days later.
  5. Isolation - a sandbox or container with the dependencies installed, no production credentials, network restricted.

What the Evidence Says

  • Benchmarks: SWE-bench Verified (resolving real GitHub issues, graded by hidden tests) and Terminal-Bench (tasks in a real terminal) are the standard measures; scores depend on the harness as much as the model, which is why leaderboards list agent systems (Evaluation and Benchmarks).
  • Task length: METR measures the length of tasks (in human time) that agents complete with 50% reliability; it roughly doubled every 7 months from 2019 to early 2025.
  • Productivity is not automatic: METR's randomised controlled trial (July 2025) with 16 experienced open-source developers working in their own large repositories found that allowing early-2025 AI tools increased completion time by 19%, although the developers expected a 24% speed-up and still believed afterwards they had been sped up. The result is specific to that setting - expert developers, mature codebases with high quality bars, early-2025 tools - but it is a warning to measure outcomes rather than impressions.
  • At the other end, OpenAI's harness-engineering report describes a team shipping a product whose code was written entirely by agents, with the humans' effort going into the environment - tests, docs, structure and review.

The common thread: output depends on the harness and the repository as much as the model, and the gains come from engineering both.

Check Yourself

Check yourself
0 / 4 answered
  1. What is AGENTS.md?
  2. A team's agent-written PRs often break conventions that reviewers catch days later. What's the most effective fix?
  3. What did METR's 2025 randomised trial find?
  4. When is a background (cloud) coding agent a better fit than an interactive one?

Exercises

Exercise - Make a repository agent-ready

Pick a repository you know. Write its AGENTS.md (under 60 lines), list the checks an agent should run and how long each takes, and identify two conventions that currently live only in reviewers' heads. Turn one into a lint rule or test.

Solution

A good AGENTS.md lists setup, test, lint and type-check commands with expected runtimes; the directory map; non-obvious conventions; forbidden actions (e.g. editing generated files, touching migrations). Converting a convention to a check - for example a test that fails when a module in api/ imports from db/ directly - turns a review comment into immediate feedback for the agent.

Exercise - Measure it yourself

Design a small study of a coding agent's effect on your own work: which tasks, how to randomise them between AI-allowed and AI-disallowed, what to measure, and how to avoid the self-report bias METR observed.

Solution

Pre-register a task list; randomly assign each task to a condition; record wall-clock time and review outcomes (defects, rework) rather than perceived speed; include review and fix-up time in the agent condition; compare with a paired or stratified analysis over enough tasks to separate the effect from task variance.

Study Notes

  • Coding agent = loop + file/search/edit/shell tools + repository context (AGENTS.md) + permissions/sandbox + verification + git
  • Modes: interactive local, background/cloud, parallel worktrees, headless CI
  • Agent-ready repositories: short AGENTS.md, fast deterministic checks, clear specs, mechanically enforced architecture, isolation
  • Evidence: benchmarks measure harness + model; METR task horizons doubling ~7 months; METR RCT found a 19% slowdown for experts in early 2025 - measure outcomes, not impressions

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท