Software Factory
Agent output now exceeds the capacity to review it, so work accumulates in front of the main branch instead of reaching it.
A coding agent produces more change than a team can evaluate. Published telemetry across millions of pull requests reports agent-assisted work arriving substantially larger, waiting far longer for a reviewer, and merging at a fraction of the rate of unassisted work, while the share merged without any review rises. The constraint has moved from writing code to deciding whether code is safe to merge. A software factory addresses that by giving agent work the structure human delivery already has: work orders sequenced by their dependencies, a harness specification per role, evidence captured as the work proceeds, and a promotion path that decides what reaches your main branch. The same structure carries into legal, financial and other regulated settings, where every automated change has to be defensible.
- Work order: dependency aware
- Planner, implementer, reviewer: a harness spec per role
- Evidence: captured per step
- Approval gate: you decide what merges
Ordered work
Work is sequenced by its dependencies and scoped to a reviewable unit.
Role specific
Planner, implementer and reviewer each get their own spec.
Gated
Policy is evaluated at five phases, and evidence decides what reaches main.
Where it applies
Teams past the pilot, where agent output now arrives faster than anyone can review it, or where audit expects a trail behind automated change.
What decides a merge
Captured evidence and a policy decision at each phase, rather than a reviewer reading a large diff under a deadline.
What you keep
Harness specifications, policy files and the promotion ledger, all checked in and diffable.
Fit
Who this suits
Teams past the pilot, where agent output now outpaces review capacity, or where internal audit expects a trail behind automated change.
Platform engineering
Asked to standardise agent output across repositories
Delivery leaders
Review capacity now caps how much ships
Regulated environments
Every automated change needs an audit trail
Maintenance backlogs
Repetitive work nobody has capacity to clear
Scope
Where agent delivery breaks, and how we address each
Each row below describes a condition that appears once agent output outpaces review. Every response runs on open source we maintain and you keep.
Where policy is evaluated
Every decision resolves to allow, ask or deny, and each one is recorded.
- 01Requestbefore work begins
- 02Responsewhat the agent proposes
- 03Tool callbefore it reaches a system
- 04Tool resultwhat came back
- 05Promotionwhat reaches main
| The problem | How we address it |
|---|---|
| Output exceeds the capacity to review it | Delivery is modelled as work orders sequenced by their dependencies, each scoped to a unit a reviewer can actually hold in their head, so throughput is bounded by the gate rather than by one person reading a large diff. |
| Merged unreviewed to beat the deadline | A promotion path moves a change through staged, canary and activated states. Nothing reaches main without passing the gate for its stage, and a rolled-back state is a first-class outcome rather than a recovery scramble. |
| One configuration for every role | Planner, implementer and reviewer each receive their own harness specification covering tools, permissions and the checks they must pass, versioned as an artifact rather than held as convention. |
| Policy written down, enforced by memory | Policy is evaluated at five points in the loop: the request, the response, each tool call, each tool result, and promotion. Every decision resolves to allow, ask or deny, and is recorded. |
| Access broader than the task requires | Work runs in a sandbox with per-role permissions and a deny-by-default posture, so the authority granted is scoped to the work order rather than to the agent. |
| No record of why a change was allowed | Evidence is captured as the work proceeds and the release decision is written to an Agent Quality Record stating what was measured, what the agent was permitted to do and who accepted the result. |
| The agent grades its own work | A harness is promoted only when a held-out task set supports it, with the split sealed from anything used to tune the harness, so a score reflects work the system has not already seen. |
| No defined path back from a failed step | Every promotion carries the version it reverts to, and the ledger records the transition, so rolling back is a documented operation rather than a judgement call at midnight. |
| No trail behind automated change | The policy decisions, the evidence and the promotion ledger together form a record that can be read months later by someone who was not present when the change was made. |
| The factory depends on who set it up | Harness specifications, policy files and the promotion ledger are checked-in artifacts. The routine is inspectable, diffable and owned by your team. |
Process
How we work
Pipeline discovery
We map how work reaches your main branch today, where review happens, what your team will let an agent merge, and where the queue is actually forming.
Workflow design
Delivery modelled as work orders with explicit dependencies, each scoped to a reviewable unit, with isolated workers and a defined path back from a failed step.
Harness and policy per role
Planner, implementer and reviewer each get a specification covering tools, permissions and checks, with policy evaluated at the request, the response, each tool call, each tool result and promotion.
Pilot and gate
One repository first, with evidence capture and the staged, canary and activated promotion path in place, then a runbook for the team inheriting it.
Deliverables
What you receive
Scope is set from your requirements rather than from a fixed package. Whatever is agreed at the outset is built into your pipeline, piloted on a repository you choose, and handed over as checked-in artifacts your team owns and can change without us.
Workflow design
Dependency-aware work orders, isolated workers, defined recovery
Role specifications
A versioned spec per role: tools, permissions, required checks
Gates and evidence
Approval points backed by captured evidence before anything merges
Runbook and handover
Operating guidance for whoever owns the pipeline after we leave
Scope
Engagement Options
Scope follows your CI, your review culture, and how much autonomy your team will grant. Cost is quoted against your pipeline once we have seen it, since the work depends on how many repositories are in scope and how far your team is willing to let a gate decide.
Typical starting point
Pilot
One repository, one workflow
Typically 6 to 10 weeks
Single pilot repository
Prove the model on contained work before the wider organisation commits.
- Pipeline discovery
- Workflow design in work orders
- Harness specification per role
- CI and headless integration
- Evidence capture and gates
- Runbook for the pilot
Programme
Multiple teams and repositories
Scoped per rollout
Custom scope
Extend a proven pipeline across teams, with governance built into the rollout.
- Everything in Pilot
- Multiple repositories and teams
- Shared specification library
- Governance and policy design
- Enablement for each team
- Phased rollout plan
Regulated
Audit and compliance first
Scoped per environment
Custom scope
Segregation of duties, retention, and documentation an assessor will accept.
- Everything in Programme
- Audit trail design
- Approval and segregation controls
- Retention and evidence policy
- Documentation for assessors
- Rollback and incident procedures
Open source
Underlying systems
Our agent-run delivery work is public. Engagements sit inside the pipeline you already have.
Questions
Questions
Discuss a Software Factory
Tell us how work reaches your main branch today and what your team is willing to automate.
London and San Francisco. Remote or on site.
