AI Harnesses

AI Harnesses

Harness engineering: how to make an AI coding agent reliable

Harness engineering shapes the environment around an AI coding agent: instructions, tests and checks. What it is, where the term comes from and how to start.

AI LLM AI Agents Harness Claude Code Software Development

Every mistake comes back

A coding agent forgets what went wrong the moment a session ends. Correct it in the prompt alone and you correct the same mistake again next week.

Put the lesson in the environment

Harness engineering turns every repeated mistake into something that lasts: a rule in the instruction file, a test, a check or a tool that makes the mistake impossible.

Work you can review

The agent hands over work that has already passed automated checks. The person reviewing it judges the substance and no longer has to catch sloppiness.

From harness to harness engineering

We wrote earlier about what an AI harness is: the software around a language model that turns it into a working agent. The loop, the tools, the context management and the guardrails. That piece explained what a harness is, and popular AI harnesses lists which ones exist. This one is about the work you put into it.

Harness engineering is shaping the environment an AI coding agent works in, so the work it delivers is work you can trust. Instruction files it reads every session, tests and checks it has to pass, tools that make common mistakes impossible, and a person who signs off on what goes live. You pick the model from a vendor. The environment you build yourself.

Where the term comes from

Harness engineering caught on quickly in early 2026. In February, Mitchell Hashimoto, co-founder of HashiCorp, described how AI found a permanent place in his work. One step he called engineering the harness: when an agent makes a mistake, take the time to build something that stops it from ever happening again.

Shortly after, OpenAI published how a small team spent five months building a product whose code was written entirely by Codex agents. The engineers typed no code themselves. They built the environment that let the agents do it.

The idea underneath is older than it sounds. Tests, linters, code review and good documentation have been around for decades. What is new is setting them up for a colleague who has forgotten everything by morning.

A new colleague who starts over every day

Think about how you onboard a new developer. You grant access to the repo, hand over a document with the team’s conventions, point to a test suite that shows what is broken, and assign someone to review the first pull requests. A few weeks in, that developer knows the conventions by heart.

A coding agent is that new developer, every session again. Fast, well-read, and with no memory of yesterday. Whatever it learned yesterday has to live in the environment. That is also where the comparison breaks down: a person remembers a correction, an agent only does once you record that correction somewhere the next session will find it.

The five layers of a good harness

1

Instructions the agent reads every session

A short file in the repo, such as AGENTS.md or CLAUDE.md, with the stack, the conventions and what is off limits. Keep it short: every line costs context, and no model reads a thousand-line instruction file well.

2

Tools that make the mistake impossible

If the agent keeps making the same mistake with a command, give it a command of its own that catches that mistake. An agent can forget a written rule. A tool that blocks the wrong path forgets nothing.

3

Checks that do not negotiate

Tests, type checks, linters and a build that has to pass. They give the same answer every time, however confident the model sounds.

4

Feedback the agent can read itself

Error messages, test output, logs and screenshots that flow back into the loop. An agent that sees the build fail fixes it before a person spends any time on it.

5

Boundaries and a person who signs off

Permissions and hooks decide what the agent may do on its own. For anything that goes live or cannot be undone, a person looks and decides.

An instruction is a request. A check is a condition. Harness engineering moves as many rules as it can from the first kind to the second.

The core rule: every mistake becomes a change to the environment

When an agent gets something wrong, the reflex is to adjust the prompt and carry on. That fixes it once. Next week, in a fresh session, it happens again.

Harness engineering asks, for every repeated mistake, what the environment was missing. The answer comes in a fixed order, from weak to strong.

1

A rule in the instruction file

The quickest fix, and the weakest. The model reads the rule, and can still drop it halfway through a long task.

2

A check that fails

A test or lint rule that turns red the moment the mistake returns. The agent sees it and fixes it within the same loop.

3

A tool or block that makes the mistake impossible

The wrong path no longer exists. This costs the most work and lasts the longest.

Pick the strongest form that is worth the effort. A mistake that happened once earns a rule at most. A mistake that returns every week earns a check or a tool.

What this looks like for us

Growthdesk, the growth system behind the site you are reading, is such a harness. AI agents work on client websites every day: content, technical fixes and ads. Here are a few things we learned along the way, each one following the rule above.

📚

Instructions per type of task

For each task the agent loads its own instruction set, with the conventions and the known pitfalls. Anything that applies to one client lives in that client's repo.

🌐

One browser command

Agents testing in the browser kept making the same handful of mistakes. Now a single command catches them, and a hook blocks the bare version. The agent no longer has to remember the pitfalls.

🏗️

No push without a green build

A push puts the change live straight away, so the build is the gate. What the agent says about its own work does not count there.

🖼️

Before and after, side by side

Every visible change comes back as a screenshot of the page before and after, with the change marked. A person judges that picture and only then approves.

🎫

Friction becomes a ticket

When an agent gets stuck on an unclear instruction or a command that behaves oddly, it files that in a backlog. People review those reports and change the harness.

There is a story behind that last point. An agent once reported that it had improved an instruction. The file was untouched, and everyone assumed the problem was solved. Since then an agent never edits its own instructions, and every change to the harness goes past a person.

What an agent says about its own work is an opinion. A green build, a passing test and a person who has seen the diff, that is evidence.

Where harness engineering goes wrong

📜

One giant instruction file

Every mistake becomes another paragraph, until the file is so long the model misses half of it. Keep the main file short and point to separate files per topic.

⚠️

Rules that should have been checks

The instruction file says: always run the tests. Run those tests in a hook or in CI, and nobody has to trust the agent to remember.

🪞

The agent approves its own work

A model reviewing its own code usually finds it fine. Let a check, a second model or a person be the judge.

🧹

Scaffolding for last year's model

Models improve fast. A workaround for a weakness that is already gone only makes the harness slower and harder to maintain. Clean up regularly.

Where to start

1

Write a short instruction file

Your stack, your conventions, how the tests run and what is off limits. One screen is enough to start with.

2

Let the agent check itself

Tests, type check and build should run with a single command, so the agent sees its mistakes before you do.

3

Keep a list of repeated mistakes

When you see the same mistake twice, turn it into a rule, a check or a tool. The order above helps you choose.

4

Put boundaries around irreversible actions

Deploys, database migrations and anything customers see: a person signs off on those. Record that in permissions, so it still holds when the agent misses the instruction.

5

Review the diff, every time

The checks catch what a machine can see. Whether the result does what was intended is for a person to judge.

That person needs something to judge against. That is why good work with agents often starts by writing down what the software should do, which is what spec-driven development is about.

Conclusion: the model brings speed, the harness brings trust

Harness engineering is the work that takes a coding agent from an impressive demo to a colleague you can rely on. The model decides how fast code appears. The environment decides whether you can trust that code. We build that environment for clients too: AI agents and automation that keep working while the models underneath change.

📌

Record every lesson

An agent forgets everything after the session. What it needs to know lives in the repo.

🚦

Checks over instructions

A failing test stops more than a rule that has to be read.

🛡️

A person signs off

Automated checks catch the sloppiness. Whether the work is right is a person's call.

Start small, with one instruction file and a build the agent can run itself. Then let the harness grow with the mistakes you actually run into.

Frequently Asked Questions

What is harness engineering?

Harness engineering is shaping the environment an AI coding agent works in, so the work the agent delivers can be trusted. That environment consists of instruction files the agent reads every session, tests and checks it has to pass, tools that make common mistakes impossible, feedback that flows back into the loop, and boundaries on what the agent may do by itself. The core rule: when the agent makes a mistake, change the environment so that mistake does not come back.

What is the difference between a harness and harness engineering?

A harness is the software around a language model that turns it into an agent: the loop, the tools, context management and guardrails. Claude Code and Codex are examples. Harness engineering is the work of tuning that harness to your codebase and your risks, with your own instructions, checks, tools and boundaries. The harness is the tool, harness engineering is how you set it up.

Where does the term harness engineering come from?

The term took off in early 2026. In February 2026, Mitchell Hashimoto, co-founder of HashiCorp, described engineering the harness as a step in how he works with AI: every mistake an agent makes is a reason to build something that stops it from happening again. Shortly after, OpenAI published a report titled harness engineering about a product a small team built entirely with Codex agents.

How does harness engineering differ from prompt engineering and context engineering?

Prompt engineering is about phrasing a single instruction well. Context engineering is about what sits in the model's context window on each turn. Harness engineering covers both and adds the rest of the environment: tests, checks, tools, permissions and the feedback loops that decide whether the agent sees its own mistakes. A better prompt helps in one session. A better harness helps in every session after it.

How do you get started with harness engineering?

Start with a short instruction file in the repo, such as AGENTS.md or CLAUDE.md, covering your stack, your conventions and how the tests run. Make sure the agent can run tests, type check and build by itself. Then track which mistakes come back and turn each repeated one into a rule, a check or a tool, from weak to strong. Put boundaries around irreversible actions and have a person review every diff before it goes live.

Do you still need harness engineering as models get better?

Yes, although what goes into it changes. Better models make some workarounds unnecessary, and you clean those up. The core stays: your codebase has conventions no model knows in advance, a deterministic check is more reliable than a model's judgment of its own work, and for irreversible actions you want a person who signs off.

AI agents that deliver work you can trust?

We build the environment around your agents: the instructions, the checks, the tools and the boundaries that make sure what goes live is right. From first setup to ongoing management.

Let's discuss your project

From AI prototypes that need to be production-ready to strategic advice, code audits, or ongoing development support. We're happy to think along about the best approach, no strings attached.

010 Coding Collective free consultation
free

Free Consultation

In 1.5 hours we discuss your project, challenges and goals. Honest advice from senior developers, no sales pitch.

1.5 hours with senior developer(s)
Analysis of your current situation
Written summary afterwards
Concrete next steps