← The Record
For CTOs, CIOs & Operating Executives

The First 90 Days of an Agent Org

Pilots die because nobody defines what graduation looks like. The agent org writes the graduation criteria on day one, then spends ninety days earning them.

Somewhere in your company, an AI pilot is having its first birthday. Nobody planned a party. The pilot still works, in the sense that the demo still demos. The steering committee still meets monthly and agrees the results are promising. Promising is the word that appears in the minutes, twelve months running, and if someone asks the only question that matters, when does this run the real thing, the room goes quiet in the specific way that means never.

The diagnosis isn't the model. The models improved all year. It isn't the team, and it isn't the budget line, which has now survived three reviews on the strength of the word promising. The diagnosis is simpler and more embarrassing. Nobody ever wrote down what graduation looks like. There is no bar the pilot could clear, because no bar exists.

A pilot with no graduation criteria isn't an experiment. It's a decision you're refusing to make, renewed monthly.

This paper is the alternative, told as a walk through the calendar. The first ninety days of an agent org, deployed inside your own perimeter, from the day it arrives holding no authority at all to the day a human hands it some. The design inverts the pilot at every step. Every stage has an exit, the exit is defined before the stage begins, and the thing being tested is never whether the demo impresses. It's whether the record earns a signature.

The graveyard has a pattern

The enterprise AI graveyard is not full of things that broke. It's full of things that never left. Deloitte's State of AI research found that only 25% of organizations have moved 40% or more of their AI experiments into production (Deloitte, 2026).

Read that number and the shape of the failure appears. These pilots were scoped as demonstrations, not as probations. A demonstration proves the thing can work. It cannot prove the thing should be trusted, because trust is a claim about behavior under authority, and a pilot, by design, holds none. So the pilot succeeds at demonstrating, and then the question "now what" arrives with no machinery to answer it. Security never gated it. Nobody defined the authority it would graduate into, so there is nothing to graduate to.

Month fourteen arrives. Promising.

The agent org runs the same ninety days in the opposite direction. It starts with the authority question, on day one, in writing.

Day one: the org interviews you

Enterprise software onboarding has one universal shape. The vendor sends forms. You fill out the forms. Weeks later, a consultant reads your answers back to you as a statement of work, and everyone pretends this was configuration.

The agent org reverses the chairs. Day one, the org interviews you.

It installs inside your perimeter, against your own model account, and it reads your repository. Then it comes back with a draft of its own constitution. Here are the files I believe I should never touch: your authentication chain, your payment logic, your migrations, the compliance code. Here are the roles I propose to staff, and what each role is scoped to reach. Here are the spend caps I propose to operate under, per agent, per day, for the whole org. Here are the rules I extracted from your contribution guide and your CI configuration, the conventions your team enforces by habit that I intend to obey by architecture.

None of it takes effect. That's the part to underline. The org can draft its constitution. It cannot ratify a word of it. Your implementer amends what the draft got wrong, and your CISO signs the registry of what must never be touched, and the signature is the CISO's, not the vendor's and not the org's. Until those signatures land, proposals are the only thing the deployment is permitted to produce.

The first thing the org ships is its own constitution, and it can't ratify a word of it.

Notice what this does to the pilot problem before the first task ever runs. The bar now exists. What the org may eventually touch, what it may spend, who approves its work, and what graduation will mean are written, reviewed, and signed while the org still holds zero authority. The pilot that dies at month fourteen never produced this artifact. The agent org produces it before lunch on day one, because the artifact is the onboarding.

Weeks one through four: everything queues

Picture week two. An agent picks up a task, drafts the change, hands it to a different agent for review, passes the build, passes the tests, passes the check that verifies no protected file was touched. In a mature deployment, that change would merge. In week two, it does something else. It stops, and it asks.

Everything stops and asks. This is probation, and it is a structural state, not a supervision mood. The full pipeline runs real work through every gate. Every action answers the five governing questions: what it touches, what it spends, who approves, what ships, what gets logged. And zero merges happen without a human click. Not few. Zero. The gate that withholds autonomy is enforced in the pipeline's code, not in a memo asking everyone to be careful, and the console wears a persistent banner saying so. The default window is two weeks of this. Hold it for four if you want. Hold it for eight. The window is displayed guidance, and the clock graduates nothing.

What you are actually doing in these weeks is reading the org's judgment before the org holds power. Every queued item arrives as a decision card: what the org wants to do, why, what it costs, what it touches. Say Tuesday brings twelve cards. Eleven are easy yeses, and the twelfth you decline, and the decline matters more than the yeses, because the reason you gave becomes part of the record, and the record is the entire point. Proposals, approvals, overrides with reasons, spend against caps, reviews survived. Probation: How an Agent Earns Authority is the paper on what that record measures and why authority should never arrive ahead of it. What's happening here is the same principle applied to a whole deployment at once.

The market has started converging on the instinct, if not the machinery: 63% of large enterprises now require human validation of agent outputs, a share that nearly tripled in a year (KPMG AI Pulse, Q1 2026). The gap between them and this is enforcement. Validation as a policy is a checkbox in a memo. Validation as probation is a gate the pipeline cannot route around.

The middle stretch: the queue thins

Picture week six. You notice something about your own yeses. You have now approved the same shape of change eleven times. Same category, same scope, same class of file, same cost band, and your answer has been identical every time.

The eleventh card isn't a decision. It's a rule you haven't written yet.

So the rule gets written. Drafting is what the org does, so the org drafts it: changes matching this shape, under this cap, touching nothing protected, no longer need a card. You ratify the rule through the same constitutional process that governs every rule, and that category leaves your queue. Forever, or until you amend it. The Ratifier Role is the paper on why this move, replacing your recurring judgment with a rule you signed, is the executive's actual job in an agent org.

This is the mechanism most people miss when they imagine ninety days of oversight, because they imagine the oversight staying constant and the humans getting tired. The opposite happens. The queue thins, and it thins for the right reason.

The decision queue doesn't shrink because you relaxed. It shrinks because your judgment got compiled.

By week ten, picture what's left in the queue: the rule amendments, the true tradeoffs, the rare item that fits no pattern because nothing like it has happened before. The routine has been ruled. The morning brief is short. And the org is now doing, in every respect but one, exactly what it will do after graduation. The one respect: it still cannot merge a line of code on its own.

Day ninety, give or take: graduation is a verb

A navy does not commission a warship because ninety days have passed since it touched water. The shipyard runs trials. An inspection board rides the ship, pushes it to full power, and writes down what it saw. Acceptance is a signature by someone accountable for the ship afterward (U.S. Navy, Board of Inspection and Survey), and a ship that later fails its material inspections can be pulled from the line. Time launches nothing and time commissions nothing. The record does, through a person.

Graduation works the same way, and it is a verb with a named subject. A human, typically the operator your CISO trusts with the room, sits down with the record the probation weeks produced. How many proposals. What share approved as-is. Every override, with its reason. Spend against every cap, every task. Review findings and what survived them. On that record, the human performs an explicit act: a confirmation, a state change, an entry in the audit trail naming who graduated the org and when. If the record isn't there, the act doesn't happen, and the org keeps queueing. Probation has no expiration date, only an exit exam.

And here is what graduation does not do. It does not retire the gates. Day ninety-one looks almost exactly like day eighty-nine. Same pipeline, same five questions answered on every action, same caps, same sacred-path registry the build still dies to protect, same kill switch sitting on the front page of the console. The only change is that a merge which passes every gate no longer waits for your click. Authority went from zero to bounded. It did not go from zero to boundless. And it stays revocable, because the state that held every merge in probation is exactly that, a state, and a state can be set back by the same class of human act that changed it. An org that earns its way out can be walked back in. No committee required, no migration, no renegotiation. One decision, one entry in the record.

What the ninety days bought

Put the two timelines side by side.

The pilot at month fourteen has a demo, a steering committee, and a slide about next steps. Ask it to prove the system should hold authority and it has nothing, because it was never structured to produce evidence, only impressions.

The agent org at day ninety has a signed constitution, a complete record of every proposal it made and every human answer it received, a spend history against caps it never breached (or a documented story about the day it tried and the gate held), and a graduation entry naming the person who read all of it and decided. Hand that to your board. Hand it to the CISO who inherits the system next year. Hand it to a regulator who asks how, exactly, an autonomous system came to hold merge authority in your enterprise. The answer is not a narrative. It's a file.

And the striking part is what the ninety days cost, which is close to nothing beyond the attention you'd owe any new hire. Nobody staffed a measurement program. The record accumulated as a byproduct of the pipeline doing real work through real gates, which is what it was going to do anyway. The pilot spends a year producing a feeling. The probation spends a quarter producing a file, from the same effort.

So go back to the pilot having its birthday down the hall. It was never one good decision away from production. It was missing the entire structure that turns operation into evidence and evidence into authority. Ninety days is not a warm-up period the vendor asks you to tolerate. It is the product working.

Ninety days isn't how long it takes you to trust the system. It's how long the system takes to earn your trust, measurably.

The takeaway: The pilot graveyard is a graduation-criteria graveyard: demonstrations scoped with no authority to graduate into. The agent org inverts the whole arc. Day one, it drafts its own constitution and humans ratify every line. For the probation weeks, the full pipeline runs and every merge queues for a human yes while the record accumulates. Through the middle stretch, your repeated approvals become ratified rules and the queue thins because judgment got encoded, not because attention wore out. Graduation is an explicit human act performed on the record, the gates all survive it, and the authority it grants stays revocable by the same kind of act. Ninety days isn't the time it takes you to trust the system. It's the time the system takes to earn it, measurably.

References

  1. SpeyAI, Rise of the Agent Org library. Companion papers: Probation: How an Agent Earns Authority (the record, and why authority never arrives ahead of it); The Ratifier Role (the human act that graduation and rule-making are instances of); Trust Scales with Structure (the series opener these ninety days put into practice).

  2. Deloitte. 2026 State of AI in the Enterprise (n=3,235). Reports that only 25% of organizations have moved 40% or more of their AI experiments into production. [CC-verified: s59 research Finding 9, adversarially verified 3-0 against Deloitte's published survey pages]

  3. KPMG. (2026). AI Pulse Survey, Q1 2026. Reports that 63% of large enterprises require human validation of agent outputs, a share that nearly tripled year over year (up from 22% in Q1 2025).

  4. United States Navy, Board of Inspection and Survey. Acceptance trials: a ship is accepted and commissioned on inspected trial performance and an accountable signature, never on elapsed time, and remains subject to material inspection afterward.

This paper is part of Rise of the Agent Org, a series by Ed Hoehn, SpeyAI. The full library is at speyai.com/record.

The missing layer

Architecture as governance. See how it runs.