Services · For engineering leaders

Coding agents that hold up on your real codebase.

Your developers have tried AI coding agents. They're impressive on a greenfield demo and unreliable on the ten-year-old system that pays the bills. The fix isn't a better model. It's a better harness.

Agentmodelharness

Everyone has access to the same models. What separates a team that ships with agents from one that cleans up after them is everything around the model: how work is specified, checked, proven and approved.

Harness engineering

What a good harness does.

Specs grounded in the codebase

Before an agent writes code, the ticket is turned into a spec that names the real files, patterns and conventions involved. Most agent failures start with a vague ticket.

Independent review

A separate reviewer with fresh context checks the work against the spec. The agent that wrote the code doesn't grade its own homework.

Evidence for every claim

"Tests pass" comes with the test output. "This is handled" comes with the line that handles it. Reviewers check evidence, not assertions.

A human merge gate

An engineer approves every merge. Agents add throughput; your team keeps the judgment and the accountability.

Team-specific skills

Your conventions, your architecture and your hard-won lessons, written down where agents can use them.

A weekly improvement loop

Every week, look at where agents went wrong and fix the harness, so the same mistake doesn't happen twice.

Open source · PRFlow

The harness, in the open.

PRFlow is an open-source harness that makes AI coding agents dependable on real, mature codebases. It's an open-source project I lead, in daily use by dozens of developers.

You can read every part of it: the spec grounding, the fresh-context review, the evidence rules, the merge gate and the self-improvement loop. Use it as is, or as a reference for your own.

PRFlow on GitHub
  • Specs grounded in the codebase
  • Independent, fresh-context review
  • Evidence for every claim
  • A human merge gate
  • A weekly self-improvement loop
The pilot

Start small enough to measure.

No big-bang rollout. A pilot proves the approach on your code, with your team, before anyone scales it.

  1. 1

    One repository

    A real one, with real history. We map its conventions and set up the harness around it.

  2. 2

    One representative ticket

    Not a toy. A ticket your team would normally pick up, run end to end through the harness.

  3. 3

    One team practice

    One habit your team adopts and keeps, like grounded specs or the weekly improvement loop.

Agent-readiness check

Is your codebase ready for AI coding agents?

Six quick questions. You'll get a readiness band and what to fix first.

How old is your main codebase?
How much would you trust your automated tests to catch a bad change?
Does every change get a real code review before merge?
How are tickets usually written?
Does CI run on every pull request?
How does your team use AI coding agents today?
0 of 6 answered

Plan an enablement pilot for your team.

Thirty minutes with Daniel. Tell me about your codebase and how your team works today; I'll tell you honestly where agents will help and what to fix first.