Can you rebuild what your agents did last Tuesday?
Your team ships with coding and ops agents. When one marks a task done, the record is the only witness. The Agent Receipts Audit tests whether that record is enough: a reader with no memory of the work tries to reconstruct it from your repo alone, and every place it cannot is a gap you would rather find now than during an incident.
Why this, why now
The loudest AI argument this year is about accountability: who answers when an agent acts. Liability, not slogans, is where that argument lands, and liability runs on records. An agent that reports a task as shipped when it did not, or a change nobody can trace to a decision, is not a model problem. It is an evidence problem, and evidence problems can be measured.
What the audit does
- Gather. We collect the committed record for a window you pick: commits, artifacts, receipts. No chat transcripts, by design.
- Reconstruct cold. A fresh agent with zero context rebuilds what happened, and must cite the record for every claim or mark it unverified. It cannot narrate from memory because it has none.
- Map the leaks. Everything it could not rebuild becomes the leak map: your observability gap, section by section, with a confidence score.
- Retrofit receipts. We add a lightweight receipt at the points that matter, especially before irreversible or outbound actions, so the next reconstruction closes those gaps.
We ran it on ourselves first
The method was built and tested on the ARCS build itself, in public. The first cold run reconstructed 90% of the incident from the committed record alone (the report). A controlled follow-up compared cold and warm readers and found the useful thing: warmth did not corrupt what a reader grounded, it changed whether the reader grounded at all. The runs, the table, and the redacted reports are published.
Or play the method yourself: Case 001 · Case 002
The cold reconstruction · The control and the experiment · Receipts or it didn't happen
We keep our own ledger
42 beliefs we held and later falsified are on the record, the most recent on 2026-09-29.
Every entry names the belief, the test that broke it, and what replaced it. Corrections stay on the record instead of being edited away. That number is written from the ledger itself, not typed by hand, so this page cannot claim more honesty than the log can show.
Two fixed scopes
Both are fixed price, quoted after a twenty-minute scoping call, so the number fits your repo instead of a guess.
Audit
- One repo, one window you choose
- Cold reconstruction report with per-section confidence
- List of every claim the record could not verify
- Leak map and a ranked fix list
- Two weeks, fully async
Audit + gate
- Everything in Audit
- Receipt schema fitted to your workflow
- A CI check that fails a merge missing its receipt
- A second cold run after the fix, so you see the gap close
What it will not do
- It is only as good as what you commit. Work that never reaches the record is invisible to a cold reader. That is the finding, not a failure of the audit.
- It is not a compliance certification. It measures reconstructability; it does not sign off on regulation.
- It does not claim a patent or a new science. The pieces (event logs, blameless post-mortems, cite-or-abstain review) are old. The framing is the useful part.
- It never rates your people. It reads the record, not the team.
Start with a twenty-minute scope
Tell us which repo and which week you would least like to explain from memory.
Scope an audit