From harness to harness engineering
We wrote earlier about what an AI harness is: the software around a language model that turns it into a working agent. The loop, the tools, the context management and the guardrails. That piece explained what a harness is, and popular AI harnesses lists which ones exist. This one is about the work you put into it.
Harness engineering is shaping the environment an AI coding agent works in, so the work it delivers is work you can trust. Instruction files it reads every session, tests and checks it has to pass, tools that make common mistakes impossible, and a person who signs off on what goes live. You pick the model from a vendor. The environment you build yourself.
Where the term comes from
Harness engineering caught on quickly in early 2026. In February, Mitchell Hashimoto, co-founder of HashiCorp, described how AI found a permanent place in his work. One step he called engineering the harness: when an agent makes a mistake, take the time to build something that stops it from ever happening again.
Shortly after, OpenAI published how a small team spent five months building a product whose code was written entirely by Codex agents. The engineers typed no code themselves. They built the environment that let the agents do it.
The idea underneath is older than it sounds. Tests, linters, code review and good documentation have been around for decades. What is new is setting them up for a colleague who has forgotten everything by morning.
A new colleague who starts over every day
Think about how you onboard a new developer. You grant access to the repo, hand over a document with the team’s conventions, point to a test suite that shows what is broken, and assign someone to review the first pull requests. A few weeks in, that developer knows the conventions by heart.
A coding agent is that new developer, every session again. Fast, well-read, and with no memory of yesterday. Whatever it learned yesterday has to live in the environment. That is also where the comparison breaks down: a person remembers a correction, an agent only does once you record that correction somewhere the next session will find it.
The five layers of a good harness
Instructions the agent reads every session
A short file in the repo, such as AGENTS.md or CLAUDE.md, with the stack, the conventions and what is off limits. Keep it short: every line costs context, and no model reads a thousand-line instruction file well.
Tools that make the mistake impossible
If the agent keeps making the same mistake with a command, give it a command of its own that catches that mistake. An agent can forget a written rule. A tool that blocks the wrong path forgets nothing.
Checks that do not negotiate
Tests, type checks, linters and a build that has to pass. They give the same answer every time, however confident the model sounds.
Feedback the agent can read itself
Error messages, test output, logs and screenshots that flow back into the loop. An agent that sees the build fail fixes it before a person spends any time on it.
Boundaries and a person who signs off
Permissions and hooks decide what the agent may do on its own. For anything that goes live or cannot be undone, a person looks and decides.
An instruction is a request. A check is a condition. Harness engineering moves as many rules as it can from the first kind to the second.
The core rule: every mistake becomes a change to the environment
When an agent gets something wrong, the reflex is to adjust the prompt and carry on. That fixes it once. Next week, in a fresh session, it happens again.
Harness engineering asks, for every repeated mistake, what the environment was missing. The answer comes in a fixed order, from weak to strong.
A rule in the instruction file
The quickest fix, and the weakest. The model reads the rule, and can still drop it halfway through a long task.
A check that fails
A test or lint rule that turns red the moment the mistake returns. The agent sees it and fixes it within the same loop.
A tool or block that makes the mistake impossible
The wrong path no longer exists. This costs the most work and lasts the longest.
Pick the strongest form that is worth the effort. A mistake that happened once earns a rule at most. A mistake that returns every week earns a check or a tool.
What this looks like for us
Growthdesk, the growth system behind the site you are reading, is such a harness. AI agents work on client websites every day: content, technical fixes and ads. Here are a few things we learned along the way, each one following the rule above.
Instructions per type of task
For each task the agent loads its own instruction set, with the conventions and the known pitfalls. Anything that applies to one client lives in that client's repo.
One browser command
Agents testing in the browser kept making the same handful of mistakes. Now a single command catches them, and a hook blocks the bare version. The agent no longer has to remember the pitfalls.
No push without a green build
A push puts the change live straight away, so the build is the gate. What the agent says about its own work does not count there.
Before and after, side by side
Every visible change comes back as a screenshot of the page before and after, with the change marked. A person judges that picture and only then approves.
Friction becomes a ticket
When an agent gets stuck on an unclear instruction or a command that behaves oddly, it files that in a backlog. People review those reports and change the harness.
There is a story behind that last point. An agent once reported that it had improved an instruction. The file was untouched, and everyone assumed the problem was solved. Since then an agent never edits its own instructions, and every change to the harness goes past a person.
What an agent says about its own work is an opinion. A green build, a passing test and a person who has seen the diff, that is evidence.
Where harness engineering goes wrong
One giant instruction file
Every mistake becomes another paragraph, until the file is so long the model misses half of it. Keep the main file short and point to separate files per topic.
Rules that should have been checks
The instruction file says: always run the tests. Run those tests in a hook or in CI, and nobody has to trust the agent to remember.
The agent approves its own work
A model reviewing its own code usually finds it fine. Let a check, a second model or a person be the judge.
Scaffolding for last year's model
Models improve fast. A workaround for a weakness that is already gone only makes the harness slower and harder to maintain. Clean up regularly.
Where to start
Write a short instruction file
Your stack, your conventions, how the tests run and what is off limits. One screen is enough to start with.
Let the agent check itself
Tests, type check and build should run with a single command, so the agent sees its mistakes before you do.
Keep a list of repeated mistakes
When you see the same mistake twice, turn it into a rule, a check or a tool. The order above helps you choose.
Put boundaries around irreversible actions
Deploys, database migrations and anything customers see: a person signs off on those. Record that in permissions, so it still holds when the agent misses the instruction.
Review the diff, every time
The checks catch what a machine can see. Whether the result does what was intended is for a person to judge.
That person needs something to judge against. That is why good work with agents often starts by writing down what the software should do, which is what spec-driven development is about.
Conclusion: the model brings speed, the harness brings trust
Harness engineering is the work that takes a coding agent from an impressive demo to a colleague you can rely on. The model decides how fast code appears. The environment decides whether you can trust that code. We build that environment for clients too: AI agents and automation that keep working while the models underneath change.
Record every lesson
An agent forgets everything after the session. What it needs to know lives in the repo.
Checks over instructions
A failing test stops more than a rule that has to be read.
A person signs off
Automated checks catch the sloppiness. Whether the work is right is a person's call.
Start small, with one instruction file and a build the agent can run itself. Then let the harness grow with the mistakes you actually run into.
What is harness engineering?
What is the difference between a harness and harness engineering?
Where does the term harness engineering come from?
How does harness engineering differ from prompt engineering and context engineering?
How do you get started with harness engineering?
Do you still need harness engineering as models get better?
AI agents that deliver work you can trust?
We build the environment around your agents: the instructions, the checks, the tools and the boundaries that make sure what goes live is right. From first setup to ongoing management.