← All articles
Explainer 9 min · September 28, 2026

What Is Harness Engineering? How AI Teams Make Agents Reliable

Harness engineering is how teams make AI agents reliable: the tools, enforced rules, verification loops and state that turn a model into a working system.

Harness engineering is the practice of designing the environment around an AI agent so it can do reliable work without constant human correction. Instead of rewriting prompts when an agent fails, engineers add what was missing: a tool, a rule the system enforces, a test that catches the mistake, or documentation the agent can actually read.

The idea caught on fast in 2026 because it matched what teams running agents in production kept finding. The model was rarely the whole problem. More often the agent was missing information, had no way to check its own output, or was free to make the same mistake twice.

Where did the term harness engineering come from?

OpenAI popularized the term in February 2026 with an engineering post by Ryan Lopopolo, Harness engineering: leveraging Codex in an agent-first world. It described a five-month experiment in which a small team built and shipped an internal product without writing any code by hand. Codex agents wrote the application, the tests, the CI configuration and the documentation. The team summed up its operating rule in four words: "Humans steer. Agents execute."

The word itself is older. Software teams have long used "test harness" for code that runs and checks other code, and Anthropic's engineering team was already writing about harnesses for long-running agents in November 2025. A month after OpenAI's post, LangChain's Vivek Trivedy gave the idea its shortest definition in The Anatomy of an Agent Harness: "Agent = Model + Harness."

What is an AI harness?

An AI harness is everything in an agent system that isn't the model. LangChain draws the line at every piece of code, configuration and execution logic outside the model itself.

A language model on its own takes in text and returns text. It can't remember yesterday's work, run code, look anything up or check whether its answer is right. The harness adds those abilities. It runs the loop that feeds results back to the model, executes the tools the model calls, stores files and memory between sessions, manages what fits in the context window, and enforces the rules the agent has to follow. Claude Code and OpenAI's Codex CLI are harnesses, and so is a custom agent a team builds directly on a model provider's SDK.

For the short version, see the glossary entry for AI Harness.

How is harness engineering different from prompt engineering and context engineering?

The three are nested. Prompt engineering shapes the instructions. Context engineering decides what information the model sees at each step. Harness engineering covers the whole system around the model, including the tools, the enforcement and the checks. LangChain describes today's harnesses as mostly a delivery mechanism for good context engineering, which is a useful way to see how the layers fit together.

Prompt engineeringContext engineeringHarness engineering
What you shapeThe instructionsWhat's in the context window at each stepThe full environment the agent runs in
Typical fix when the agent failsReword the promptAdd, trim or restructure informationAdd a tool, rule, test or feedback loop
ScopeOne model callOne task or sessionEvery run, across sessions

The difference shows up when something breaks. A prompt fix asks the model to try harder. A harness fix changes the environment so the mistake can't happen again, or gets caught automatically when it does.

How does harness engineering work in practice?

Teams building serious agents have landed on a similar set of practices. None of them is exotic. Most are ordinary engineering discipline applied to a new kind of worker.

Give the agent a map instead of a manual

OpenAI's team first tried putting everything the agent needed into one large AGENTS.md instruction file. It failed. The file crowded out the actual task, went stale quickly and made every rule look equally important. They replaced it with a file of roughly 100 lines that works as a table of contents, pointing the agent to a structured docs folder that serves as the system of record.

The broader lesson is that an agent only knows what it can see in its context. A decision made in a Slack thread or a meeting is invisible to it until someone writes it down where the agent looks.

Enforce rules in code, not in prose

Written guidance gets skimmed or misread. Rules that run automatically don't. OpenAI's team enforced its architecture with custom linters and structural tests, and wrote the linter error messages so they tell the agent how to fix the problem. Once a rule is encoded that way, it applies to every line the agent writes.

Build loops that verify the work

An agent that can't observe the results of its work has to guess whether it succeeded. Good harnesses give it ways to check: test suites, a browser it can drive, logs and metrics it can query. OpenAI wired Chrome DevTools and a per-task observability stack into the agent's environment so it could reproduce bugs and confirm fixes on its own.

There's a catch. Anthropic found that agents asked to grade their own work tend to praise it, even when the quality is mediocre. Its fix was to separate the agent doing the work from a second agent that evaluates it, tuned to be skeptical.

Carry state across long tasks

Long jobs outgrow a single context window. Anthropic's long-running harness uses an initializer agent that sets up the project, writes a feature list and a startup script, and hands off to a coding agent. That agent completes one feature per session, commits to git and leaves a progress note for the next session. The files do the remembering, so each fresh session picks up where the last one stopped.

Clean up drift continuously

Agents copy the patterns they find, including bad ones. OpenAI's team initially spent every Friday, about 20% of the week, cleaning up low-quality agent output. That didn't scale. They encoded their standards as mechanical rules in the repository and set background agents to scan for violations and open small refactoring pull requests on a regular schedule.

Treat every failure as a missing piece

This is the mindset underneath the rest. When an agent struggled, OpenAI's team asked what capability was missing and how to make it visible and enforceable for the agent. The answer might be a tool, a rule, a test or a document. Whatever it is, it goes into the harness so every future run benefits.

What difference does a harness actually make?

With the same model, a better harness can be the difference between a demo and something that works. The clearest public numbers come from the labs and tool builders doing this work, so read them as self-reported. They all point the same way.

Anthropic ran one prompt, a request for a 2D retro game maker, through Claude Opus 4.5 twice: once as a single agent, and once inside a three-agent harness with a planner, a generator and an evaluator (source).

SetupRun timeCostResult
Single agent20 minutes$9Looked plausible, but the core game didn't respond to input
Full harness6 hours$200Richer editors and a game that actually played

The harness cost more than 20 times as much. It also produced the version that worked.

LangChain reported moving its coding agent from the top 30 into the top 5 on the Terminal Bench 2.0 leaderboard by changing only the harness.

OpenAI's experiment produced roughly a million lines of code and about 1,500 merged pull requests over five months. Three engineers started the project, averaging 3.5 pull requests per engineer per day, and the team estimated it built the product in about a tenth of the time hand-coding would have taken. OpenAI is careful to add that the results depend on that repository's specific structure and tooling, and shouldn't be assumed to carry over without similar investment.

Does harness engineering apply beyond coding agents?

Yes. Coding is where the term took hold because code is easy to test automatically, but any agent doing real work needs the same parts. A claims-intake agent needs tools with the right permissions, business rules it can't override, checks on its output, a record of what it has already done and a way to hand off to a person when judgment is required.

For a business, the harness is also the durable part of the system. Models get replaced every few months. The harness holds your business rules and data connections, and it keeps working when you swap in a better model. That's why Custom AI Studio builds the harness as code the client owns.

Will better models make harness engineering unnecessary?

Some of today's harness will disappear, but the discipline won't. Anthropic's framing is that every harness component encodes an assumption about what the model can't do on its own, and those assumptions go stale as models improve.

Anthropic watched that happen within a few months. An earlier harness needed full context resets because Claude Sonnet 4.5 tended to wrap up work early as its context filled. Opus 4.5 largely stopped doing that, so the resets came out. When Opus 4.6 arrived, Anthropic removed the sprint structure too, while keeping the planner and evaluator because they still earned their place. The conclusion was that the space of useful harness designs moves rather than shrinks as models get better.

The habit that follows: when a new model ships, re-test your harness. Remove the pieces that no longer carry weight, and look for tasks that are newly within reach.

Frequently asked questions

What is harness engineering in simple terms? Harness engineering is building the environment an AI agent works in so it gets things right more often. That means giving it the right tools and information, enforcing rules automatically and checking its work.

Who coined the term harness engineering? OpenAI popularized it in a February 2026 post by Ryan Lopopolo. The word "harness" was already in use: software teams have long built test harnesses, and Anthropic wrote about agent harnesses in November 2025.

What is the difference between an agent harness and an agent framework? A framework is a toolkit you use to build a harness. LangChain or the Claude Agent SDK give you building blocks, and the harness is the specific system you assemble from them for your agent, with its own tools, rules, memory and checks.

Is Claude Code a harness? Yes. Claude Code and OpenAI's Codex CLI are both agent harnesses wrapped around a model. LangChain notes that products like these are post-trained with the model and harness in the loop together, which is part of why they work well out of the box.

Do I need harness engineering if I use an off-the-shelf agent? Yes, in a lighter form. Instruction files like AGENTS.md, skills, hooks, permissions and test commands all configure the harness. Tuning them is harness engineering on a small scale.

Written by

Emi Yakushev

Emi Yakushev is a Product Marketing Specialist at Custom AI Studio, where she runs content and SEO and writes the studio's case studies and explainers on agentic AI, AI agents, and custom AI builds. Previously a marketing strategist at Zenna Consulting Group.

Read more.

All articles →