ai

What Is an Agent Harness?

The software around an LLM that turns it into an agent: the loop, tools, context, and guardrails. Where the term came from and why more isn't better.

An agent harness is the software wrapped around a language model that turns it into an agent: the loop that calls the model repeatedly, the tools it can execute, the context management that decides what the model sees, and the guardrails that decide what it’s allowed to do. The model predicts text. The harness is everything that makes those predictions add up to work getting done.

Anthropic’s Claude Code documentation puts it in one line: “Claude Code is the harness; Claude is the model inside it.” OpenAI uses the same word for the shared execution core behind every Codex surface (CLI, IDE extension, web), all of them powered by the same Codex harness. When both labs independently settle on a term, it’s worth pinning down what it means.

One thing it doesn’t mean: Harness.io, the CI/CD and software delivery company. Same word, coincidental collision, unrelated domain. If you searched “harness devops” and landed here expecting deployment pipelines, that’s them. This post is about the LLM concept, though if you build deployment pipelines for a living, stick around: you already have most of the mental model.

Fig 1The agent loop
Model
Predicts text. It can ask for a tool. It cannot run one.
“I need to see which tests fail.”
① intent →
bash("npm test")
← result ③
3 failed · 41 passed
HarnessALLOW
$ npm test
✕ auth.spec.ts  3 failed
✓ 41 passed      2.4s
1 gather context→ 2 call model→ 3 run tool→ 4 check result↻ Turn 1 / 5
The model can only ask. The harness does.Each turn, the harness runs the requested tool, checks it against policy, and sends the result back, until the tests pass or a limit stops it.

Model, Scaffold, Harness, Agent

These four words get used interchangeably, and they shouldn’t be. A Hugging Face glossary essay from May 2026, itself a sign the terminology had gotten muddy enough to need one, draws the lines this way:

  • Model: the LLM. It can express intent to call a tool. It cannot execute anything.
  • Scaffold: the behavioral configuration. System prompt, tool descriptions, how responses get parsed, what carries over between steps.
  • Harness: the execution layer. It calls the model, runs the tool calls, feeds results back, and decides when to stop.
  • Agent: model + harness. Something that acts, not just responds.
Fig 2Four layers
Agent
Harness
looptoolscontextpermissionshooks
Scaffold system prompt · tool specs · parsing
Model
Layer 3 of 4
Harness

The execution layer. It calls the model, runs the tool calls, feeds the results back, and decides when to stop.

Examples

Claude Code · Codex · Aider · mini-swe-agent

The model sits inside the scaffold, which sits inside the harness.Click a layer, or use the buttons, to see where one ends and the next begins.

In casual usage “harness,” “scaffold,” and “framework” blur together, and for a blog post that’s usually fine. The distinction that actually matters is model versus everything-else, because the everything-else is where most of the engineering lives — and, as we’ll see, where a surprising amount of benchmark performance comes from.

Where the Term Came From

“Harness” is old software vocabulary. A test harness is code that sets up conditions, drives the thing under test, and scores the output. LLM research inherited that sense directly: EleutherAI’s lm-evaluation-harness, started in the GPT-3 era, became the standard way to benchmark models and the backend of Hugging Face’s Open LLM Leaderboard. In that world the harness was deliberately boring: standardized scaffolding, so that what you measured was the model, not an accident of prompt engineering.

Fig 3From test rig to runtime
  1. 2020EleutherAI's lm-evaluation-harness becomes the standard way to benchmark models.
  2. Sep 2025Anthropic renames the Claude Code SDK to the Claude Agent SDK.
  3. Feb 2026“Harness engineering” gets a name, from Mitchell Hashimoto and OpenAI.
  4. May 2026Hugging Face publishes a glossary to settle the words.
A measuring deviceThe runtime for an agent
The same word, two jobs.First the harness measured models. Once models could use tools, it became the thing that runs them.

Then models learned to use tools, and the word migrated. Once an LLM is taking actions in a real environment, the code that drives it stops being a measurement device and becomes a runtime. A datable marker of the shift: in September 2025, Anthropic renamed the Claude Code SDK to the Claude Agent SDK, on the reasoning that “the agent harness that powers Claude Code can power many other types of agents, too.” By early 2026 the term was everywhere. Mitchell Hashimoto’s February 2026 post crystallized “harness engineering” as a named practice, and OpenAI published an essay with that exact title the same month, describing a team that shipped a product with zero manually-written code. Their framing of the new engineering job: “design environments, specify intent, and build feedback loops.”

A fact-check aside, since I nearly repeated this myself: the term is often attributed to Andrej Karpathy. His widely-cited 2025 year in review doesn’t contain the word “harness” at all. The vocabulary came out of the labs and the eval community, not a single coinage.

What’s Actually Inside a Harness

Every serious harness is built around the same loop (gather context, take action, verify the result, repeat), but the implementations differ in revealing ways. A quick tour of the ones I use or have studied:

Fig 4Where each part acts
  1. 01context
  2. 02model
  3. 03policy
  4. 04pre hook
  5. 05run tool
  6. 06post hook
Hooks

Scripts that fire at fixed points in the loop: before a tool runs, after an edit, at session start. The model never decides whether they run, so hooks are the part of the harness you control fully.

.claude/settings.json
"hooks": {
  "PreToolUse": [{
    "matcher": "Bash",
    "hooks": [{
      "type": "command",
      "command": "./rewrite.sh"
    }]
  }]
}
Every harness runs the same loop. The parts differ in where they act on it.Pick a part to light up the stations it controls.

The loop and tools. Claude Code ships file access, shell execution, and search as first-class tools, plus subagents: child instances with their own context window and restricted tool access, so exploration doesn’t pollute the main conversation. Codex runs one shared Rust core under every product surface.

Context management. The context window is the scarcest resource, and each harness spends it differently. Aider builds a “repo map” — a graph-ranked summary of your codebase’s important symbols, compressed into a token budget, sent with every request. Claude Code compacts: when the window fills, older tool outputs get cleared and the conversation gets summarized. Cursor assembles context from a workspace index plus rules files that activate by file-glob.

Memory files. Nearly every harness converged on the same idea: a Markdown file in your repo that gets injected at session start. Claude Code reads CLAUDE.md; Codex, Cursor, and twenty-plus other tools read AGENTS.md, now an open format under the Linux Foundation. The harnesses differ; the convention is shared.

Permissions and guardrails. This is where harnesses look most like infrastructure. Claude Code layers permission rules (deny → ask → allow) over sandboxed shell execution. It’s IAM policy thinking applied to a model’s tool calls.

Hooks. Deterministic scripts that fire at lifecycle points: before a tool runs, after an edit, at session start. On my machine, a hook rewrites git and other CLI calls through a token-optimizing proxy before they execute, and another injects a reminder to persist session learnings into my local memory store. The model never decides whether those run. That determinism is the point: hooks are the part of the harness you control completely.

If that list reads like a platform engineering backlog (isolation, resource budgets, policy, lifecycle events, observability), that’s not an accident. I’ve argued before that harness engineering is a DevOps skill; this is the anatomy behind that claim.

More Harness Isn’t Better

Here’s the part that surprised me. Given how much engineering goes into these harnesses, you’d expect the elaborate ones to decisively beat simple scaffolds. The measured answer is: not reliably.

METR tested this directly in February 2026, running the same models under production harnesses and under deliberately simple scaffolds. Claude Code against bare-bones ReAct (an agent that just takes an action, sees the result, and repeats) was a statistical coin flip: Claude Code won in 50.7% of bootstrap samples. Codex against METR’s generic Triframe scaffold actually lost most of the time, winning only 14.5% of samples. And mini-swe-agent, a harness in roughly 100 lines of Python, scores above 74% on SWE-bench Verified — competitive with systems orders of magnitude more complex.

Fig 5Production harness vs simple scaffold
Claude Code beats a bare ReAct loop50.7%
Codex beats METR's generic Triframe scaffold14.5%
mini-swe-agent, about 100 lines, on SWE-bench Verified74%+
The elaborate harness did not reliably win.Top two bars: share of bootstrap samples in which the production harness won. Bottom bar: task success rate. Sources: METR (February 2026) and the SWE-agent project.

So the harness doesn’t matter? No, the opposite. Swapping scaffolds changes what the same model scores, which is exactly why METR controls for it when measuring capability. What the elaborate harness buys you just isn’t raw benchmark points. It’s everything a benchmark doesn’t measure:

  • Safety: permission gates, sandboxes, and cost caps that make it survivable to let an agent run unattended. A 100-line loop with full shell access benchmarks fine right up until it doesn’t.
  • Ergonomics: memory files, hooks, and skills that encode your project’s conventions, so you stop re-explaining them every session.
  • Recoverability: compaction, session resumption, observable tool traces. The difference between an agent you can debug and one you re-run and hope.

One line from a Hacker News thread on harness engineering sums up the practitioner view: “A decent model with a great harness beats a great model with a bad harness.” The benchmark data says the sophistication isn’t free capability. The lived experience says it’s what makes the capability usable. Both are true, and the tension between them is basically the design brief for every harness team right now.

That brief keeps expanding. Anthropic’s latest iteration lets Claude generate its own orchestration harness per task — the harness stops being a fixed artifact a human designs once and becomes something the agent composes on the fly. Whether that’s the future or a detour, it tells you where the labs think the payoff is.

Why This Matters to You

If you’re choosing between coding agents, you’re mostly choosing between harnesses. The frontier models are closer to each other than the scaffolding around them is. Compare them on harness terms: how they manage context, what their permission model lets you safely automate, whether their memory and hooks let you encode your conventions once.

And if you build one, even a script that collects CI failure logs, asks a model what broke, and posts the answer to Slack, you’re doing harness engineering.

The model is the part you rent. The harness is the part you own.

That’s where your effort compounds. For how to approach building them with the infrastructure skills you already have, see Harness Engineering: The DevOps Skill Nobody Told You About.