A concept note

Ratchet: building complete software systems with AI agents

An agent can build the software. So what is the agency for? This sets out where the value actually sits once the typing is automated — and how Ratchet is built to capture it.

For  Developers & the business Subject  Orchestrating agents into whole systems Status  Concept note — for evaluation
Section 01

The problem

AI models can now write code, and the agents built on them can do more than write. They can read a codebase, run commands, use tools, and keep working until a task is finished. For a software development agency this raises a direct question: if an agent can build the software, what is the agency for?

The answer matters, because it is easy to assume the value was in producing code. It was not. What follows sets out where the value actually sits, and how Ratchet is built to capture it.

Section 02

The landscape

It helps to be clear about what these tools are and are not.

A model on its own reasons but cannot act. An agent wraps the model so that it can: it works in a loop, reads and writes files, runs commands, and uses whatever tools it is given. Within a single task, and when the intent is clear, an agent is very capable.

Two things follow from this.

That last point is the agency's real work. An agency turns what a business needs into software that works, and it is paid for domain knowledge, judgement, and the ability to spot risk — not for typing. Agents change how that work is done. They do not remove the need for it.

The opportunity

Take everything an agency knows and use it to direct an agent, systematically enough to build a whole system rather than a single task. That is what Ratchet does.

Section 03 / the core

Ratchet

Ratchet is an orchestrator. It takes a system that exists only as documented intent and directs an agent, across many runs, until that system exists as working software.

It has three parts.

Part 01

Design corpus

The documented intent. Vision, requirements, constraints, architecture, roadmap, chosen design approaches — the documents that say what is being built, why, and what must be true of it. The agency's domain knowledge and judgement, written down in a form the process can act on.

Where senior effort goes
Part 02

Phasing

The work is broken into ordered steps that build on one another — each small enough for the agent to do well, and clear enough for a person to follow.

Ordered, buildable steps
Part 03

Evaluation loop

Each phase is executed, then checked against defined criteria. Pass and it is accepted; fall short and it is refined and run again. What this is — more than running an agent and hoping for the best.

The trust mechanism

The design corpus comes first: everything downstream is only as good as this corpus, which is where the senior people spend their effort. Phasing turns it into a sequence the agent can execute one step at a time. And the evaluation loop is what makes Ratchet more than running an agent and hoping for the best.

The evaluation loop — per phase
RUN
Execute
The agent does the phase.
→
CHECK
Evaluate
Measured against defined criteria.
→
PASS
Accept & settle
Work moves on. Settled work is not unpicked later.
↶ Falls short → refine the phase and run it again, until it passes.

Work that has passed is settled and is not unpicked later, so the system builds up steadily rather than drifting.

First dimension / specialised, deterministic tooling

Give the agent tools that don't improvise

Two further ideas sit alongside the core. The first is about the tools the agent is given to use.

An agent is probabilistic: ask it the same thing twice and you can get two reasonable answers, one of them subtly wrong. For much of a software system that variation is a liability — data access, API scaffolding, anything with a correct shape where deviation is simply risk.

The answer

Give the agent specialised tools that are deterministic — producing the same correct output every time from an explicit definition. MakerGen is an example: it generates code from model definitions, identically on every run. Where the agent would otherwise improvise such code, it instead calls a tool that guarantees it.

This changes how the model is used. Rather than asking the agent to do everything, it coordinates specialised tools and spends its own reasoning only where reasoning genuinely helps. Quality goes up, variation goes down, and engineering standards are built in rather than hoped for.

Each such tool is high-leverage: it captures a standard once and applies it on every project afterwards.

Second dimension / the process trail and its documents

What the process leaves behind

The second idea is about what the process leaves behind. A system built this way produces more than code. The process is document-centric, and it leaves a durable record of what happened.

Each phase is bracketed by documents.

Before

Planning document

States what the phase will do and how, before any code is written — a brief for the agent and a checkpoint for the human.

→
During

The phase runs

The agent executes against the brief; the evaluation loop holds it to the criteria.

→
After

Evaluation document

Records how the result measured against its criteria, what was found, and what was changed in response.

Read in sequence, these turn the build into a history that can be followed. A developer joining late, a client asking why something was done a certain way, or a reviewer auditing the work can read the decisions and their reasons rather than reconstruct them from the code. The reasoning survives, not just the result. And because these documents are reusable, the agency's judgement carries forward from one project to the next instead of being lost when the code ships.

Alongside the documents, the process records itself — what ran, in what order, how long it took, and what it cost. This is not added on for invoicing; it is a direct output of running the work in an orchestrated way.

Why it matters commercially

The result is work that can be measured: cost per phase, cost per deliverable, time to completion. That makes the whole model commercially legible. It can be estimated, priced, and accounted for against real figures rather than guesswork.

Section 06 / who stays in command

The developer's role

It would be a mistake to read all this as the developer doing less. The opposite is true. As more of the routine execution is handled by the agent and its tools, what remains for the person is the harder part: deciding what is correct, setting the direction, and carrying responsibility for the result.

Automation of this kind does not lower the bar for the developer. It raises it.

The developer is the expert who guides the process and stays accountable for what it produces. That begins before any code is generated, in authoring the design corpus — turning a business's needs into intent precise enough for the process to act on. It depends on knowing what good looks like, in the domain and in the engineering, because the quality of everything the agent produces is bounded by the quality of what it is given.

During execution the developer sets the criteria each phase is judged against, reads the results critically, and decides what to accept and what to send back. This is where experience earns its place. An agent's output is often plausible and occasionally wrong in ways that are not obvious, and catching that requires someone who already knows the right answer well enough to recognise the wrong one. The developer also knows when to step in directly — when a problem is better solved by hand than handed back to the loop.

Across all of it the developer is coordinating parts that have to fit together: the documented intent, the specialised tools, the phased work, and the evaluation that holds it to standard. Building and maintaining those tools is part of the role, as is judging which parts of a system should be made deterministic and which are better left to the agent's reasoning.

The skills this calls for

Not the skills of writing code quickly. They are the ability to frame intent clearly, deep domain and engineering knowledge, sound judgement about quality and risk, and the discipline to run a structured process and hold to it. The work is less about producing each line and more about knowing what should be produced, recognising when it has been, and remaining responsible for the whole. The process runs the execution; the developer remains in command of it.

Section 07 / the conclusion

Where the value sits

Put together, the agency's role does not shrink. It shifts toward higher-leverage work.

Package the knowledge

Domain knowledge written into a design corpus the process can act on.

Build the tools

Deterministic tools that hold the system to a known standard.

Orchestrate the work

Direct the agent through phased execution and evaluation.

Because the process documents and accounts for itself as it runs, all of this comes with a record attached — the decisions, the reasons, the time, and the cost — which is what lets the agency show its working to a client and put a defensible number on it.

None of this is less skilled, and all of it is more leveraged. The same expertise that once produced a single system now produces the corpus, the tools, and the orchestration that can produce many — faster, more consistently, and with less rework.

The agent supplies capability. The value lies in the knowledge fed into it, the tools it is given to use, the discipline of the process it runs inside, and the record it leaves behind. That is what Ratchet is for.
Where this fits

The concept of the line. The map of the tooling. The orchestrator that runs it.

Ratchet is the orchestration layer of the DARK LINE. See the line as a whole, or how it maps to tooling already in flight.

See the line →

Or read the automation map →