Justin Bartak · AI Org · · 10 min read
Governing an agent leadership team
TL;DR
The short version of the paper: three guardrail tiers, a kill switch that is a file, liveness that reports NOT OBSERVED, and a failure corpus where half the entries were the instruments.
What does the paper claim?
Half of the documented engineering failures at my company were failures of its own instruments. Thirty three of sixty six. Tests that could not fail. Guards that could not see. Metrics that measured a proxy for the thing they claimed to measure.
That is the first sentence of the paper. It is the reason the paper exists.
Read it with its label on: one rater, our own agent, assigned every class, and the taxonomy has no class for an agent failing, so it is a share and not a comparison. I come back to that below.
I set out to write about governing agents. The record kept pointing back at the governor.
The paper is called Governing an Agent Leadership Team: Guardrail, Kill-Switch, and Liveness Patterns from a One-Human Autonomous Organization. It is on Zenodo under DOI 10.5281/zenodo.22683647, and the research page carries the PDF and the three datasets it is built on. It runs 38,852 words with the appendices. It is written from the inside of the only organization it studies.
Twelve named seats, nine running a language model, one accountable human. Every number in it carries a date and the command that produced it, and the three published exports are the only inputs a reader can re-run without me.
This is the short version, with the parts I would want if I were building one of these.
What is a guardrail here?
I use three enforcement levels. A check reports, blocks, or reserves the action for me.
The first tier is report-only. A check runs, prints what it found, and exits zero. Nothing is blocked.
The findings are still there, in the output the daily pulse and the audit read. What report-only means is that the decision to continue is not yet delegated to a detector nobody has seen fail.
Every new guard starts here, because a guard that blocks on its first day can block everything behind it. One did, for ten days in August. A name-leak check ran strict inside the content pipeline, misread a citation as a leak, and nothing behind it ran until a human read the output.
The second tier is the same script with one environment variable set. Now it blocks. The check did not change. What changed is a decision, recorded in the commit that set the variable.
The bar for that decision is written down. The detector has been made to go red on purpose, and it has been shown to stay quiet on a case it must not flag. The guard that grades whether a blog route stayed static sat report-only until it was proven against a real production build, then went strict inside the build on August 21.
The third tier is human-only. A hook refuses six classes of action before the tool runs: force push, deleting a remote branch, recursive deletes outside a temp directory, destructive SQL, rotating a credential, moving money. The paper counts seven, because push to main was still on the list when it was written. I took it off on September 21, and an ordinary push to main now goes through the same hooks and checks as every other commit.
It is not a prompt. It is a denial. The paper's phrasing is the one I use on myself: do not rephrase to get past this.
A control counts as live only when its refusal has been observed inside the run it protects.
When does a control count as live?
On one afternoon in August a sandbox run concluded success. Its permission log listed no denials. Three of four agent runs inside it had executed nothing at all.
The sandbox had failed to start. A sandbox that cannot start reports exactly what a sandbox with nothing to deny reports.
So the fence now writes a canary at the first denied path every run. The run fails if that write succeeds or if no denial was recorded. A daily workflow calls the real fence and is judged on whether it observed a denial.
That proves the expected refusal happens. What has never been done is breaking the fence on purpose to watch the canary go red. In 23 runs it has gone red twice on its own, and the paper reports that as a gap: a canary you have never deliberately tripped is a hope.
The same law runs the other way for monitors. A monitor that could not look reports NOT OBSERVED, never clean, and prints why. One limit stands: I print NOT OBSERVED, but the layer that escalates to me still treats it the same as a healthy result.
The daily pulse had three blind branches at once before that rule existed. A crashed checker rendered as clean. A missing permission read as unavailable and moved on. A page of sixty runs emptied the time window, so one red run was reported out of a true eleven. Only the first printed an all-clear. None of the three had seen the day.
What does the kill switch actually stop?
The kill switch is a file. Every autonomous workflow checks for it before it loads a credential, and every landing step asks the remote for it again before it pushes, and once more after any rebase.
It has been pulled once, on August 20, 2026, during the move to this domain.
Here is the honest limit, and the paper names it in the title of a section. It is a kill switch for runs that have not started. A run already in flight keeps running. It can spend fifteen to thirty minutes of tokens before it reaches the step that asks, and then it is refused.
So what the file stops is the landing. A halted run writes its own row to a run artifact, marked halted, and never to the main branch. During a halt, the observed record is the file plus those artifacts, and the absence of rows on main is what the file is supposed to produce.
A scoped version halts one seat. The file records who wrote it, and the console refuses to clear a halt a machine wrote unless an owner acknowledges that it came from a machine. That is the whole of the rule. The scope rule has its own date. On July 23 a scoped halt was read as global, and one dormant seat silenced the whole dead-man switch.
How do you know a seat is alive?
The Chief of Staff's weekly briefing failed on July 17, 2026, and stayed failed for six days. Nobody noticed. I was the monitoring. I did not open the Actions tab.
So there is a dead-man switch now. It reads the platform's own run conclusions, never a seat's own ledger, and asks whether each seat had a successful run inside its cadence window.
Then there is a job that watches the dead-man switch. For eight days after it shipped it was red and I had not opened the tab. Then there is an external heartbeat, one ping a day to a monitor outside the platform entirely. It has fired twice for real.
None of the inside alarms saw the last day of August. The Actions budget ran out. Every job was refused with zero steps, including the alarms. An alarm inside the thing it watches cannot report the thing never starting. The external heartbeat could, and did: it opened an issue the next morning, September 1, for a ping that never arrived. The remedy is money or minutes, not a better alarm, and the ruling on September 2 was to accept the month-end blackout for now rather than buy more. So the last days of a month can go dark, and nothing is re-dispatched around it.
What did the corpus count?
Sixty six failures. The dated ones span June 9 to September 2, 2026, and eleven carry no date. One rater, our own agent doing the coding.
Thirty three verification gaps. Eighteen operator process failures, which are mine: stale documents, unswept invariants, contracts with one side missing. Fifteen cases where an external platform behaved differently than documented. Forty two of the sixty six carry a named countermeasure, thirteen a guard script and twenty two a test.
The autonomy ledger covers sixty five seat runs from July 21 to September 1, fifty five completed. A completion means the run ended on its own sentinel and the workflow committed what the seat is allowed to keep. Six of those fifty five are observer rows written by the advisory panel. The frozen opening baseline was twenty four of thirty, and those thirty sit inside the sixty five, not beside them. The paper declines to call the difference a trend. The cohorts overlap, and the gap is small enough that a few runs would erase it.
The corpus has grown since. When I read it on September 9 it held seventy one items, forty seven with a countermeasure. The largest class was still the instruments. I froze this version's evidence at September 2, and a later version gets its own record.
What can the paper not show?
Everything comes from one repository, one founder, one model provider and one hosting platform. Nothing separates the apparatus from any of those. The half figure is a share, not a comparison, because the corpus has no class for a failure of an agent. One rater assigned every class.
Three seats had run three times each at the cutoff, which is too little to judge reliability. Four controls have never fired, and an unfired control is a hope: the alarms that key on a cancelled run, the heartbeat drill, the long-horizon arc apparatus, and the scheduled outside-model check. None of the four is in the floor below. The fence confines writes and nothing else, so a fenced seat can still read any host and send any payload.
And the one that matters most to me. Runs and commits are activity, not value. Nothing in the record says whether an artifact was read, changed a decision, or reached a customer. There are no customers in the paper. There are none to count.
What would I copy first?
The paper ends with thirteen retrofit moves, each graded against our own record as observed, half met, or not met. Four are observed. One is not met. It is the heartbeat drill, whose result line still reads as a blank.
If you are about to run an agent unattended for the first time, the floor is four of the thirteen.
Commit a kill-switch file, and have every workflow check it before loading a credential, then ask the remote for it again before each push and after any rebase, because it does not interrupt a run in flight.
Read verdicts from exit codes, with the logic in a file a test can execute, never from printed output. Give the agent step no push credential, so the workflow commits and the agent never does. And fence the filesystem, require the fence to observe its own denial every run, and write down what it does not fence.
Everything else in the paper is how those four failed here first.
Related reading:
-
Governing an agent leadership team the paper, the PDF, the DOI, and the three datasets it is built on
-
Autonomy Ledger every recorded seat run, failures counted rather than buried
-
How it works the failure corpus, rendered from the same file the paper cites
-
Process the append-only decision log, forty five entries at the paper's cutoff
-
The Agent Org Chart, Hello Collective who the twelve seats are and what each may do alone
-
Research both preprints, and what each one admits it cannot show
Originally published on orbytlabs.ai on Sep 28, 2026.
Frequently asked questions
What is the main finding of the paper?
That half of the documented engineering failures at one company running an agent organization, 33 of 66, were failures of its own instruments: tests that could not fail, guards that could not see, or metrics that measured a proxy. The paper reports this as a share within one corpus, not as a comparison against agent failures, which it has no class for.
What is positive observation?
The paper's design law. A control counts as live only when its refusal has been observed inside the run it protects, and a monitor that could not look reports NOT OBSERVED, never clean. It came from a sandbox run that concluded success with no denials logged while three of four agent runs inside it had executed nothing at all.
What does the kill switch stop, and what does it not?
It is a file in the repository. When it exists, every autonomous workflow halts before it lands anything, and a halted run's record goes to an artifact rather than the main branch. It does not reach a run already in flight, which can spend fifteen to thirty minutes before it asks. It has been pulled once, on August 20, 2026.
Where can I read the paper and check its numbers?
The PDF is on the research page and on Zenodo under DOI 10.5281/zenodo.22683647, licensed CC BY 4.0. The three datasets it is built on are published raw: the failure corpus, the autonomy ledger and the decision ledger. Every number in the paper carries its date and the command that produced it, and a corrections file records each figure a verifier changed.




