What a Shipped Change Costs
Every AI vendor sells outcomes nobody can price. A governed agent org prints the number: one shipped change, all-in, denominated in millionths of a dollar.
It is the third Thursday of the quarter and the CFO is walking the budget, line by line, asking every owner the same question. What does one unit of the thing you make cost?
Marketing answers in four seconds: cost per acquired customer, trending down. Support knows the cost of a resolved ticket to the dime, and knows which ticket type is wrecking the average. The plant, if there is a plant, has known its cost per unit since before anyone in the room was born.
Then the walk reaches the AI line. It is one of the bigger lines now. And the answer to the same question is a demo, an internal satisfaction survey, and a slide titled "Momentum."
Nobody in the room can price one outcome. Not roughly. Not at all. The line got funded on a story, renewed on a story, expanded on a story. And the CFO, whose entire job is converting stories into numbers, lets this one line live on faith.
AI is the only line on the budget where the spend is exact and the value is an anecdote.
This paper is about the moment that stops being true. A governed agent org can answer the CFO's question with a number. Not a modeled number, not a survey-weighted productivity estimate. A number, for one unit of outcome: what it cost to ship that change, all-in. Every token, every review, every retry.
And the reason it can answer is not a clever reporting feature. The reason is the governance. Hold that thought. It is the second half of the paper.
Exact spend, anecdotal value
Here is the strange part. Half of the AI line is the most precisely measured spend in the company. Vendors price by the seat, by the token, by the call. Nothing on the input side is vague. You know what you paid down to the API call, because the vendor made sure of it.
What nobody knows is what you got per dollar. KPMG's pulse survey has US enterprises above a billion in revenue averaging $207 million in AI budgets over the next twelve months (KPMG, 2026). Nine-figure budgets, committed against an outcome nobody in the room can isolate well enough to price. Immaculate input pricing. No unit of output.
That combination has a name in every other procurement category.
If you cannot price one outcome, you did not buy outcomes. You bought activity.
The vendors are not lying, exactly. They sell outcomes in the deck and price inputs on the invoice, and the gap between those two documents is not malice. It is that they cannot invoice an outcome, because in their architecture an outcome has no edges. The work is smeared across chat sessions, autocomplete acceptances, and the human cleanup nobody logs. When Gartner forecast that over 40% of agentic AI projects would be canceled by the end of 2027, the cited reasons were escalating costs and unclear business value (Gartner, 2025). "Unclear business value" is the analyst's polite phrase for a room full of people who could not price one outcome.
The unit is the shipped change
So pick a unit. Not the token, which is an input. Not the conversation, which is activity. Not the story point, which is a negotiable fiction that inflates every January. The unit that survives contact with a finance team is the shipped change: a discrete piece of work that cleared review and landed in production. Engineering already thinks in it. The business can see it.
A governed agent org can total the cost of that unit because of two facts of construction.
The ledger meters everything, as it happens. Every model call is written to a ledger at the moment it occurs: which agent, which task, what it cost, denominated in millionths of a dollar. Not floats, not estimates trued up at month-end. When the Agent Spends Money made the control argument for this ledger, the caps checked in the execution path before a call is made, the halt when a cap trips. This paper is about the other job the same ledger turns out to do.
Every piece of work is bounded before it starts. A shipped change begins its life as a Brief: a written scope with acceptance criteria and a cost band, ratified by a human before any agent moves. It ends at a merge gate, where a human clicks. Everything between those two points is attributable to it. Including the failures. The implementation attempt that got rejected in review and redone is in the total for that change, because it was work the change required. The ledger does not launder the retry into overhead. It bills it to the change that needed it.
Bounded work plus a metered ledger means the division is trivial. Total attributable cost, over one shipped change. Then the rollups: the median this month, the trend, the cost by task type, the expensive tail and which agent is producing it. Each of those is a query, not a study.
The other column
Now the comparison, because a unit cost with nothing next to it is trivia.
Here is a worked illustration, and I want to be plain that it is an illustration, not telemetry. Say a mid-level engineer costs $130,000 in salary. Load it the way your finance team already loads it, benefits, payroll taxes, the recruiter's fee amortized, the laptop, the seat, and you land somewhere near $180,000 a year, call it $90 an hour. Now price a modest change through a human pipeline: half a day of engineering, an hour of a reviewer who had other plans, a slice of a standup, a ticket shepherded across a sprint. You are at roughly $1,000 before anything has gone wrong, and something usually goes wrong.
Run the same change through the governed org, again as illustration. Say the intake interrogation that produced the Brief cost forty cents of model spend. The decomposition, a dollar. The implementation, three dollars across two attempts, the failed one included on purpose. The independent review, ninety cents. Say the whole thing lands near six dollars, plus one minute of human attention at the merge gate.
The point is not six dollars against a thousand. Pick your own numbers; the shape survives. The point is that one of those columns is a measurement, and the other has never actually been computed. Estimated in planning meetings, yes. Computed, no.
The product ships this comparison as a first-class surface, and the honesty rules around it matter more than the ratio. The human benchmark is a labeled assumption, seeded at $1,000 of fully loaded cost per merged change and editable by the customer, because a benchmark you cannot argue with is a benchmark you should not trust. The frame is marginal model cost against fully loaded human cost, with the operator retained as the approver: a new instance holds every merge for a signature, and after graduation code still reaches a person unless the operator deliberately opens the small-change window. It is a throughput comparison, not a headcount claim. And keeping the human at the gate is not a hedge we invented for comfort: 63% of large enterprises now require human validation of agent outputs, roughly triple the prior year (KPMG, 2026). The market is converging on the gate. The question is who can price what flows through it.
The number is a by-product of the governance
Here is the part that boards keep missing. You cannot bolt cost-per-outcome onto an ungoverned system. The hard part is not arithmetic. The hard part is attribution, and attribution is the one thing governance cannot help producing.
Walk the chain. The cost band exists because the org refuses to start work without a ratified scope. The review record exists because no agent is allowed to approve its own work. The retry is attributable because every task runs in its own bounded session against its own branch. The merge is timestamped and owned because a human gate demanded it. The log line exists because logging is itself a gate, not a courtesy. None of those artifacts was created for accounting. Every one of them was created to control the org. Stack them up, and an audit trail of authority turns out to be an itemization of cost.
The governance was installed to constrain the work. It accidentally invoiced it.
Cost-per-outcome is not a reporting feature. It is a property of an architecture in which no work happens off the books.
Try to produce the same number from a loosely-run copilot deployment and you hit the wall immediately. Where does one outcome begin? Which of the four hundred autocomplete acceptances belong to it? How much of the senior engineer's Thursday was cleanup, and against what do you book it? The work has no edges, so the cost has no home. You cannot total what you cannot bound, and you cannot bound what nothing ever forced to have a boundary. The vendors selling "productivity uplift" are not withholding the unit number from you. They do not have it, and their architecture cannot make it.
This is why the deeper argument of this series keeps landing in the same place. The Anatomy of a Governed Task walked one task through every gate. Logs Are Not Audit argued that a queryable record beats a pile of log files. This paper is those arguments with a currency symbol attached: every gate you install for safety turns out to mint the evidence for value.
One ledger, two jobs
The finance instinct is to treat controls and reporting as separate systems, a thing that stops bad spend and a thing that describes spend afterward. In a governed agent org they are one object. The same ledger row that the cap checks before a model call is the row the cost-per-change query reads afterward. The control and the proof are not reconciled with each other, because they were never two things.
The same ledger that caps the spend is the one that proves the value.
That claim cashes out in a document. The org generates a progress report for the person above the operator, the one who approved the budget and will never log in to anything. Work shipped and in flight against the ratified Briefs. Spend against budget for the period. What the gates caught. Value delivered against the labeled benchmark. The report is compiled under an extraction-only discipline: every figure must trace to a value in the authoritative store, and a figure that is not there renders as missing rather than being invented. Honest zeros stay zeros.
That discipline is the opposite of how AI progress is currently reported upward. The system is not allowed to flatter you. A report that cannot fabricate a number is a report a board can lean on, and it comes from the same place the kill switch lives.
The question changes
I run SpeyAI on the agent org it sells. When I want this number for my own company, I do not commission an analysis. I query it. Our own instance's cost per shipped change for the most recent period is $1.44. Illustrations aside, that sentence is the difference in kind. Not that the number is small. That the number exists, on demand, with the ledger rows attached.
So here is the board-level reframe, and it fits in one sentence. Stop asking "is the AI initiative working," a question that can only ever be answered with a story, and start asking "what does one shipped change cost, and which way is the number moving." The first question produces the Momentum slide. The second produces a unit cost, a trend line, and an argument you can win or lose with arithmetic.
The vendors who cannot answer it will call the question unfair. It is not unfair. It is the same question every other line on the budget answered years ago, and the AI line has been excused from it for exactly as long as nobody in the room could see an alternative. There is an alternative. It looks like governance, because it is governance, and the number falls out of it.
The faith-based era of the AI line ends the first time someone in the room can print the unit cost. Be the one holding it.
The takeaway: Every AI vendor prices inputs and sells outcomes: the work has no edges, so the cost has no home. A governed agent org prices the outcome itself. Every model call lands in a ledger denominated in millionths of a dollar, every piece of work is bounded by a ratified Brief and a human merge gate, and the division falls out: cost per shipped change, all-in, retries included, next to a labeled, editable benchmark for the fully loaded human cost of the same unit. The number is only possible because governance made every step an artifact, and the same ledger that caps the spend is the one that proves the value. Change the board question from "is it working" to "what does a shipped change cost, and which way is it moving."
References
SpeyAI agent org architecture. Live reference at speyai.com. Per-call spend metering to a micro-USD ledger with caps enforced in the execution path, cost bands ratified on the Brief before work begins with actuals reported against them, cost-per-shipped-change reporting against a config-seeded fully loaded human benchmark with a labeled-assumption caption, and a generated no-login executive report compiled under an extraction-only rule (every figure traces to the authoritative store or renders as missing).
KPMG. (2026). AI Pulse Survey, Q1 2026. US enterprises above $1 billion in revenue averaging $207 million in AI budgets over the next twelve months; 63% of large enterprises now requiring human validation of agent outputs, nearly three times the prior year (up from 22% in Q1 2025).
Gartner. (2025, June 25). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Press release citing escalating costs, unclear business value, and inadequate risk controls as the primary failure modes.
This paper is part of Rise of the Agent Org, a series by Ed Hoehn, SpeyAI. The full library is at speyai.com/record.