← All insights
Insights · Engineering · 5 min read

Agent = model + harness

Every team has access to the same AI models. What separates teams that ship with coding agents from teams that clean up after them is the harness around the model.

Same models, very different results

Talk to a few engineering teams about AI coding agents and you’ll hear two completely different stories.

One team says agents are now part of how they ship. Agents pick up real tickets, the work gets reviewed like anyone else’s, and throughput is up without quality going down.

Another team, often with a similar codebase and the same models, says agents are a mess. They write plausible code that breaks conventions nobody wrote down, claim tests pass when they don’t, and generate more review work than they save.

The difference usually isn’t the model. Both teams have access to the same ones. The difference is everything around the model. I find it useful to say it as an equation:

Agent = model + harness

The model is the part everyone talks about. The harness is the part that decides whether you can trust the output.

What a harness is

A harness is the set of practices, tools and rules that wrap a model when it does real work. It answers questions the model can’t answer for itself:

  • What exactly is this change supposed to do, in this codebase?
  • Who checks the work, and with what context?
  • How do we know a claim about the work is true?
  • Who decides it’s ready to merge?
  • How do we stop the same mistake from happening next week?

On a greenfield demo, you can skip most of these. The codebase is small, the conventions are obvious, and nobody is depending on the result. On a mature system with years of history, skipping them is exactly why agents fail.

Five parts of a good harness

1. Specs grounded in the codebase

Most agent failures I see start before any code is written, with a vague ticket. “Add export to the reports page” means something very specific to the engineer who’s worked in that code for three years. To an agent, it’s an invitation to guess.

A grounded spec turns the ticket into something concrete: which files are involved, which existing patterns to follow, which edge cases matter, and what “done” means. It’s written against the real code, not from memory.

2. Independent review

The agent that wrote the code shouldn’t be the one that decides it’s correct. A separate reviewer, with fresh context and the spec in hand, checks the change. Fresh context matters: a reviewer that inherited the author’s assumptions will miss the same things the author missed.

3. Evidence for every claim

Agents are fluent, and fluency is persuasive. “All tests pass” and “this case is handled” read as true whether they are or not.

A good harness requires evidence for every claim: the actual test output, the line of code that handles the case, the query result. Reviewers, human or AI, check the evidence, not the assertion. This one rule removes a surprising amount of wasted review time.

4. A human merge gate

An engineer approves every merge. Agents add throughput; your team keeps the judgment and the accountability. This isn’t a temporary safety measure to drop once the agents get better. It’s how you keep ownership of your own codebase.

5. A weekly improvement loop

Every week, look at where agents went wrong, and fix the harness rather than the individual change. Add the missing convention to the team’s written skills. Tighten the spec template. Add the test that would have caught it. Over time, the harness accumulates your team’s hard-won knowledge, and the same mistake stops recurring.

Why this matters more on mature codebases

Mature codebases are where most of the business value lives, and they’re also where agents struggle most. They carry undocumented conventions, historical decisions, and fragile areas that experienced engineers route around without thinking.

A harness makes that implicit knowledge explicit, in a form agents can use. That’s also, not by accident, good for the humans. New engineers onboard faster into a codebase with grounded specs and written conventions.

PRFlow: a harness in the open

PRFlow is an open-source harness that puts these ideas into practice: specs grounded in the codebase, independent fresh-context review, evidence for every claim, a human merge gate, and a weekly self-improvement loop. It’s an open-source project I lead, in daily use by dozens of developers. You can read every part of it on GitHub, use it as is, or use it as a reference for your own.

How to start

Don’t start with a big rollout. Start with one repository, one representative ticket, and one team practice. Run that ticket end to end through a harness, look honestly at what went wrong, and fix the harness. Then do it again.

If you want a quick read on where your team stands, the agent-readiness check asks six questions about your codebase and practices and tells you what to fix first.

Free resource

The AI Readiness Scorecard

Twenty questions across data, process, people and governance. Score your company and see where to focus first. Delivered by email as a web page and a printable version.

Want to talk about how this applies to your business?

Thirty minutes with Daniel. No pitch deck: you talk about your business, I tell you honestly where AI fits and where it doesn't.