Contents
Map

18 · Agent Engineering

Overview

View as:

18 - Agent Engineering

Agent engineering is designing the system around the model: the harness it runs in, the loops and graphs that structure its work, the context it sees, the verifiers that decide whether it succeeded, the skills and memory that let it reuse know-how, and the specifications it works from. This module draws on the published practice of teams building coding and long-running agents, examines coding and computer-use agents as case studies, and ends with a lab where you build a harness and measure one of its components.

Learning objectives 8-10 hours (notes + lab)
By the end of this module you will be able to:
  • Design an agent harness - state across context windows, verification in the harness, context management - and ablate it as models improve
  • Decide when to build a loop, when to split it into a graph, and what it must iterate against
  • Explain how coding and computer-use agents work, set a repository up for agents, and read the productivity evidence critically
  • Package procedures as SKILL.md skills and design memory policies
  • Build verifiers and environments that resist reward hacking, and relate them to agent RL
  • Run spec-driven development with review at the points that matter

Where This Module Fits

flowchart LR
    H["01 Harness"] --> LG["02 Loops &<br/>graphs"]
    H --> CX["03 Context"]
    CX --> CA["04 Coding<br/>agents"]
    CX --> CU["05 Computer use"]
    H --> SK["06 Skills &<br/>memory"]
    LG --> VE["07 Verifiers, environments,<br/>agent RL"]
    CA --> SD["08 Spec-driven<br/>development"]
    VE -.-> LAB["🧪 Lab: minimal<br/>harness"]
    H -.-> LAB

    style H fill:#d8dfe8,stroke:#b0bac8
    style VE fill:#e8e0d4,stroke:#c8b89a
    style LAB fill:#dde4dc,stroke:#b0c4b0

Chapter Map

#ChapterYou will learnTime
1Harness EngineeringAgent = model + harness; state across context windows; verification in the harness; ablation50 min
2Loops and GraphsWhen a loop pays off; the verifier as bottleneck; graph signals; the "dies halfway" test45 min
3Context Engineering for AgentsContext rot; just-in-time context; clearing, compaction, offloading, notes, sub-agents45 min
4Coding AgentsAnatomy, modes of use, agent-ready repositories, benchmark and productivity evidence45 min
5Computer-Use and Browser AgentsPerceive-act loop, API vs DOM vs pixels, environment safeguards, injection risk40 min
6Skills and MemorySKILL.md and progressive disclosure; where knowledge belongs; memory policies45 min
7Verifiers, Environments and Agent RLVerifier types and rubrics; environments; RLVR for agents; reward hacking50 min
8Spec-Driven DevelopmentVibe coding vs vibe engineering; specify → plan → tasks; where humans review35 min
9Q&A Review Bank30 questions across the module45 min

Code Lab

LabWhat you buildRuns on
A Minimal Coding HarnessA coding-agent harness with file/edit/shell tools on eight repository tasks graded by hidden tests; measures a verification stop hookA local OpenAI-compatible model (laptop) or any hosted API

Mini-Project

Take one repository you work in and make it agent-ready, with evidence:

  1. Write its AGENTS.md and one skill for a recurring procedure.
  2. Build ten tasks from real past issues, each with hidden tests.
  3. Run a coding agent (the lab harness or a product) on them before and after your changes, three trials each.
  4. Report pass rate, claimed-but-failed rate and cost, and which change mattered - with the traces that show why.

Review

  • Q&A Review Bank - consolidated questions for this module
  • Module quiz - every Check Yourself question in this module, in course order

Previous: 17 - Production Agents · Next: 19 - Solutions Architecture & Communication

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.

⚡AI-assisted content - always verify, always explore multiple perspectives·