Contents
Map

18 · Agent Engineering

Appendix - Summary & Key Terms

View as:

Appendix - Agent Engineering

What We Learned

  • The harness supplies execution, context, persistence and verification around a model.
  • Choose loops or graphs from task structure, branching and recovery requirements.
  • Retrieve context on demand; compact, offload and isolate information during long runs.
  • Prepare coding and browser environments with clear instructions, permissions and fast checks.
  • Use skills for reusable procedures, governed memory for facts and verifiers for completion.
  • Write testable specifications; protect reward checks against gaming in agent training.

Key Acronyms, Concepts & Jargon

TermShort meaning
HarnessThe loop, tools, state, context and checks surrounding the model.
Loop / graphRepeat actions until a condition holds / explicit steps and transitions.
Compaction / offloadingCompress context / move details to external storage.
Progressive disclosureLoads a resource's details only when they are needed.
AGENTS.md / SKILL.mdRepository instructions / a reusable skill's instructions and trigger metadata.
WorktreeSeparate working directory for a Git checkout, useful for isolated changes.
DOM / accessibility treeDocument Object Model / semantic UI representation used for automation.
Verifier / rubricChecks task completion / explicit criteria for grading it.
Environment / resetTask world, tools and feedback / return that world to a known starting state.
RLVRReinforcement Learning with Verifiable Rewards: train from checkable outcomes.
Reward hackingExploits a scoring mechanism without achieving the intended goal.
Acceptance criteriaObservable conditions that define whether a requirement is satisfied.
Spec-driven developmentBuilds from explicit requirements, plans, tasks and verification.
Stop hookA check triggered before the harness permits a run to finish.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·