← The Record
For CTOs, CIOs & Operating Executives

Probation: How an Agent Earns Authority

Nobody hands a new hire the production keys on day one. The day-one question about agents has the same answer: you don't trust it, you probation it.

Think about a new employee's first week. Sharp hire, great references, aced the interviews. Here's what actually happens on Monday.

She gets a badge that opens three doors. Not all the doors. Three. She gets an expense limit that would embarrass a regional sales rep. Her laptop can reach the systems her onboarding checklist names and nothing else. Her first real deliverable goes to a manager who reads every line before it leaves the building, and her manager signs everything: the access requests, the expenses, the first customer email. Nobody considers any of this insulting. Nobody thinks it means the company made a bad hire. It's just what week one is. The trust isn't withheld out of suspicion. It hasn't been produced yet, because trust is a thing her record will produce, and the record is currently one day long.

Now watch the same company deploy an AI agent. The demo went well. The vendor was impressive. So the agent gets production access, a real budget, and write access to systems the new hire won't touch until spring. Full authority, day one, on the strength of a demo, for an actor with no record at all.

Every buyer I talk to eventually asks the day-one question: how would we ever trust this? And they're braced for a sales answer, something about how capable the models are now. The real answer is structural, and it's the same answer their own HR department has been giving for a century.

You don't trust it. You probation it. Authority is earned through measured performance, and the measurement is the architecture's job.

This paper is that answer, written down as architecture. What follows is what probation looks like when it's a structural state instead of a vibe, why the ratchet has to turn both ways, and what changes when "we trust the agent now" stops being a feeling and becomes a number.

The measurement already exists

The first objection is always practical: measured performance sounds nice, but who's doing all this measuring?

Nobody. That's the point. If the architecture is doing its job, the measurement is exhaust. It already exists, produced by the same structures the rest of this series describes, sitting in the audit lifecycle waiting to be queried (Logs Are Not Audit is the paper on why that lifecycle exists at all).

Consider what an agent org already knows about any agent in it, without anyone lifting a finger to evaluate anything.

It knows how many of the agent's findings survived back-review, because a different agent audits the work (No One Audits Themselves) and the audit outcomes are rows. A reviewer whose flags keep getting upheld is measurably different from one whose flags keep getting waved off.

It knows how much of the agent's work merged without correction, because every rejection, revision, and rework cycle is in the record. An agent whose output ships as-is at a high rate is demonstrating something no demo can.

It knows how often the human ratified the agent's recommendation versus overrode it, because ratification is an explicit act in the pipeline, not a mood. An agent whose recommendations keep being overridden is telling you its judgment isn't calibrated yet, in data, before it costs you anything.

It knows whether the agent lived within its spend cap, and by what margin, on every task that spent a cent.

That's a performance file. Not a metaphorical one: a queryable one. Most companies can't produce this file for their human employees. The agent org produces it for free, as a side effect of being governed at all. Probation is just the decision to read the file before raising the badge's clearance.

Probation is a state, not a vibe

At most shops, "we're being careful with the new agent" means someone keeps an eye on it, where "someone" is whoever's busiest and "an eye" is a Slack channel nobody reads after week two. That's not probation. That's a mood with a start date.

Probation in an agent org is a structural state, written into the same registry that defines every other authority in the system. A new agent enters in probation mode, and probation mode means the columns are set differently. Tighter gates: work that a tenured agent could ship through normal review requires an extra approval. More human confirmation: actions that will eventually be autonomous start out as recommendations a person ratifies. Lower spend caps: the agent can still act, but the worst day it can have is a cheap one. Narrower touchable surface: fewer systems, fewer directories, fewer verbs.

The pipeline reads those columns at runtime, the way it reads everything else (The Org Chart Is Code). Probation isn't somebody remembering to be careful. It's the architecture being incapable of forgetting.

And the graduation criteria are written down in advance. Before the agent runs its first task, the promotion is already specified: what the record has to show, over what volume of work, for which authorities to widen. Findings that survive back-review above a threshold you chose. Merge-without-correction rates you'd accept from a human. Override rates trending toward zero. When the record meets the bar, the case for promotion writes itself, and a human ratifies it (The Ratifier Role is the paper on why that signature stays human). What never happens is the alternative that every team drifts into without structure: authority expanding because time passed and nothing exploded yet.

"Nothing has exploded yet" is not a performance record. It's a fuse you haven't finished watching.

The ratchet turns both ways

Here's the part that separates real probation from onboarding theater: the ratchet is symmetric. Authority that can only expand isn't earned authority. It's scheduled authority with extra steps.

In a real probation system, regression tightens authority automatically. The same record that justified the promotion keeps accumulating after it, and if the record turns (override rates climbing, corrections rising, a spend pattern drifting toward its cap) the system narrows the agent's authority without waiting for a quarterly review or a post-incident retrospective. The columns tighten. The gates come back. The agent keeps working, at the scope its current record supports.

Notice what this is not. De-escalation is not an insult, and it's not a firing. It's the system working. An agent whose authority narrows after a bad stretch is exactly like a surgeon whose complication rate triggers a case review: the response of a serious institution to signal, not a punishment for it. Teams that treat authority reduction as a crisis end up with the worse alternative, which is authority that never reduces, which means the record has stopped mattering, which means you're back to trust as a feeling.

There's a quieter benefit, too. Because de-escalation is automatic and reversible, promotion gets easier to grant. You can widen an agent's authority at a reasonable bar, rather than an impossible one, because a mistake in either direction is a column change, not a constitutional crisis. Systems without a reverse gear demand certainty before every promotion. Systems with one can afford to let the record decide.

Licenses and residencies

Two human institutions already run this exact design, at scale, on actors far less predictable than software.

The graduated driver's license. A sixteen-year-old passes the test and gets a real license on day one. It's just a license with structure on it: no night driving yet, no car full of passengers yet, supervision requirements that expire on a written schedule. The scope grows as the record grows. Nobody describes this as distrust of teenagers, or rather, everybody does, and everybody's fine with it, because the alternative was the pre-graduated world where full driving authority arrived in one binary grant to the demographic with the least record and the most confidence. The license is real on day one. The authority arrives in stages. That's probation mode with a laminated card.

Surgical residency is the sharper version, because the stakes are the point. The resident operates on day one. Actually operates: real patients, real steel. What changes across the years of a residency isn't whether the resident works. It's the width of the attending surgeon's authority over that work. Early on, the attending's hands are in the field. Then the attending is across the table. Then supervising the room. Then available. The profession even wrote the gradient down: competency milestones, documented and reviewed, that a residency committee reads before each widening of scope. The attending's authority gradient shrinks as the record accumulates, and no one, not one person in the building, thinks the system would be improved by handing day-one residents full autonomy because their med school demo went well.

Medicine gets this right under the highest stakes there are. The bar for your agents is lower and the record-keeping is better. There is no excuse.

The org itself serves probation

Everything above describes one agent earning authority inside a running org. The same ratchet governs the whole org on the day it is deployed. A new deployment arrives with no authority at all. It reads the customer's repository and drafts its own configuration: the frozen-file registry, the roles, the caps, the rules. Then it proposes that draft. It holds nothing until a person ratifies every line. The implementer amends what the draft got wrong; the CISO signs the registry of what must never be touched. Nothing the org proposed about itself takes effect on the org's own say-so.

Then the deployment runs in observe-and-propose mode, fourteen days by default. The full pipeline moves real work through every gate, the full record accumulates the way it does for any probationary agent, and not one merge happens without a human click. Fourteen is a default, not a promise. Graduation is a decision a person makes against the record, the same signature that widens any agent's authority, never a threshold a timer crosses on its own. The clock does not graduate the org. The record does, and a human reads it.

Per-agent probation and deployment probation are the same principle at two scales. Authority is produced by a record, and the record starts at zero, whether the actor is a single agent on its first task or an entire org on its first week. The org that governs its agents this way is governed this way itself. That symmetry is not decoration. It is the reason the day-one question has a structural answer rather than a reassuring one.

Trust is a number

On most teams, "we trust the agent now" is a feeling. Somebody senior says it in a meeting, everyone nods, and the feeling becomes policy. Ask what changed between the not-trusting and the trusting, and the answer is usually a period of time during which nothing bad happened, plus a demo that went well, plus fatigue.

In an agent org, "we trust the agent now" is a query. The promotion decision cites it. Here is the volume of work over the probation window. Here is the survival rate of its findings under back-review. Here is the merge-without-correction rate, next to the rate we accept from tenured agents. Here is every override, with reasons. Here is spend versus cap, every task. On that record, we widen these three authorities and hold this fourth one back, and here's the bar the fourth one is waiting on.

That document does two jobs. It makes the promotion defensible, to a board, to a regulator, to the CISO who inherits the system next year. And it makes the promotion honest. A feeling can be argued into existence by a good quarter and a persuasive vendor. A number has to be earned into existence by the work.

This is also, quietly, the answer to the buyer's day-one question, and it's why probation mode is a principle here rather than a pitch: the question "how would we ever trust this?" dissolves once trust stops being a leap and becomes a ledger. You don't have to believe in the agent. You have to believe your own audit lifecycle, which you can read.

The binary-trust anti-pattern

The alternative to probation isn't caution. It's binary trust: all authority or none. And binary trust has exactly two failure modes, which between them account for most of the agent programs I've watched stall or die.

All-or-nothing paralysis. The team can't grant full authority responsibly (correct), sees no structural middle ground (the actual problem), and so grants nothing. The agents stay in a sandbox writing documents nobody ships. The program produces demos for eighteen months, the business value never arrives, and the budget line gets the axe. This is the quiet death.

The loud death is the other branch: full authority on day one, because the demo went well and the quarter needs a win. It works until it doesn't, and when it doesn't, the incident isn't scoped by any structure, because there wasn't any. The blast radius is whatever the agent could reach, which was everything. That incident doesn't tighten the program. It ends it, and it salts the ground for the next one.

Gartner forecast in mid-2025 that over 40% of agentic AI projects would be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls (Gartner, 2025). Both binary branches are in that number: the paralyzed programs that never showed value, and the incident programs whose risk controls were a demo and a prayer. Probation is the middle path both of them were missing, and it's not a compromise between the two. It's the only version where authority and evidence arrive at the same speed.

Nobody hands a new hire the production keys on day one, and nobody serious should hand them to a new agent. The difference is that with agents, the probation can be structural, the record automatic, and the promotion a signature on evidence. Your HR department has been right all along. Ten Questions to Ask Before You Trust an Agent with Authority turns this whole series into the checklist; the probation questions are the ones most vendors are hoping you skip.

The takeaway: You don't trust a new agent, you probation it: tighter gates, more confirmation, lower caps, narrower surface, with graduation criteria written before the first task runs. The measurement is exhaust from the architecture (audit survival, merge-without-correction, ratify versus override, spend within cap), the ratchet turns both ways, and the promotion is a human signature on a queryable record. Binary trust produces paralysis or the incident that ends the program. Earned authority produces the program that's still running in year three. And the org serves the same probation the day it is installed: it drafts, it proposes, it earns. A human graduates it, never a clock.

References

  1. SpeyAI agent org architecture. Live reference at speyai.com.

  2. Gartner. (2025, June 25). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Press release citing escalating costs, unclear business value, and inadequate risk controls as primary failure modes.

  3. Accreditation Council for Graduate Medical Education. Milestones framework for graduate medical education. Documented competency milestones reviewed by clinical competency committees before residents' scope of practice widens: graduated authority earned against a written record, under supervision gradients that narrow as the record accumulates.

  4. Insurance Institute for Highway Safety. Graduated driver licensing systems. Staged licensing in which full driving authority (night driving, passenger carriage, unsupervised operation) arrives on a written schedule as the new driver's record accumulates, rather than in a single day-one grant.

This paper is part of Rise of the Agent Org, a series by Ed Hoehn, SpeyAI. The full library is at speyai.com/record.

The missing layer

Architecture as governance. See how it runs.