9/10/2026

From Coding Assistants to a Governed Delivery System

What we learned building an AI Software Factory — and what the measurements actually show

From Coding Assistants to a Governed Delivery System
AUTHOR
Mauro Krikorian
-
Head of R&D
https://www.linkedin.com/in/maurok/
Southworks AI Software Factory in a Nutshell

I wrote this white paper, not as a market commentary or a prediction — but as a first-hand account of an ongoing effort my team and I started back in 2025, when we set out to answer a question our customers kept raising in different words: AI was making individual developers faster, so why was delivery not getting faster with them? Every architectural decision, failure, and open gap in it comes from our own delivery work across real client engagements. Half a year later, I think we have an answer worth publishing — along with the architecture, the tradeoffs, and the numbers behind it.

‍

The first wave solved a narrow problem

Coding assistants help people produce code faster. Enterprise delivery is larger than that. Requirements have to be interpreted, repositories have to evolve together, approvals have to be captured, failures have to be recovered from — and every step has to be visible and auditable. Every tool we tried could write code. None could hold state across a wait, resume after a failure, or carry a lesson into the next work item. Those turned out to be the delivery lifecycle, not the periphery.

One principle organizes everything we built: AI agents should reason, and workflows should coordinate. Workflows own state, approvals, retries, and recovery. Agents handle planning, implementation, and problem solving inside controlled execution environments. That separation is what made the system reliable. Human involvement is what made it adoptable.

‍

The shift is economic as much as technical

A development center sells hours, and the buyer carries the risk that those hours produce less than expected. A delivery line inverts that: the unit being bought is a reviewed, merged change, and its cost is a measurement rather than a forecast. On standard, recurring work — roughly two-thirds of a typical backlog — that measurement is an order of magnitude in our favor on speed, and considerably more than that on cost.

I resisted the temptation to publish a single headline multiplier. The paper gives a full chapter to it instead: what each ratio is, which baseline it is measured against, the population it was computed over, why speed and cost figures should never be multiplied together, and why the honest whole-backlog number is smaller than the per-feature one. It also covers what happens to review capacity and throughput over the first quarters — and names the one dimension we have not measured yet.

‍

The parts most people underestimate

Governance is what unlocked adoption, not what slowed it down. Full autonomy was never the goal. Approval is a workflow state rather than a side conversation: a plan can be approved, sent back, or escalated, and a reviewer's comment becomes a remediation input rather than a note someone has to act on by hand. Automated quality gates run before any person is involved, which is what keeps reviewer attention on judgment instead of on defects a machine should have caught.

Failure is a route, not a full stop. A failed check or a requested change is classified and routed — implementation-level fixes back to implementation, design-level problems back to planning — which is what makes imperfect first-pass output still worth having.

Learning is the part that compounds. The feedback humans give at the gates becomes reusable, reviewed guidance in a version-controlled knowledge base. A team improves as its individuals learn, so that improvement lasts as long as their tenure. A delivery line improves by writing down what it learned, so the guidance outlives the project, the engineers, and the review thread it was said in.

Where it runs is part of the design. The control plane, the execution sandboxes, the datastores, the telemetry, and the knowledge base all sit inside the client's own tenancy. Nothing is routed to a vendor-hosted service — which also means none of the numbers in the paper depend on us to be believed. They are read from audit tables in the client's environment, and the organizations that produced them can re-derive any of them without our involvement.

‍

We also wrote down what we cannot yet prove

Planning still varies across runs on identical work, and closing that gap is our first priority. We can route any stage to any approved model, but we cannot argue with complete certainty which model belongs where — providers, versions, and models keep changing underneath a running system, so the right choice is not a fixed answer but one that depends on the scenario, the context, the enterprise, and the cost budget it has to fit. And everything we measured is speed, cost, and volume — a defect-rate comparison against a conventional baseline is the measure that would settle the quality question outright, and it is the next thing we intend to put beyond dispute. The paper states all of that plainly, along with the risks we manage, because a model that only describes what works is of little use to anyone deciding whether to adopt it.

‍

What is in the full paper

·       The three grades of AI adoption we passed through, and the ceiling each one hit.

·       Six architecture phases, including the two decisions that shaped everything downstream — and the ones we had to reverse.

·       The anatomy of a delivery line, its eleven stages, and where it deliberately stops.

·       What it actually takes to stand one up: eight setup activities, and the two most often compressed.

·       The full measurement chapter — definitions, baselines, populations, and what the ratios do not cover.

·       What happened when we deliberately killed a running process, fifteen lessons learned, and a seven-step guide to where a team should start on Monday.

If there is one line to take away, it is this: AI-enabled software engineering is an orchestration discipline before it is a generation capability. The gains we can measure came from the orchestration, governance, and learning built around the agents — not from the agents alone.

‍

Download the full white paper — Building an AI Software Factory: From Coding Assistants to Governed Delivery Systems.