Notes from the Trenches Verdict: found, fixed, receipted

A validation question, and what it exhumed

The Question That Found the Outage

It began as the smallest possible ask, how do I know the new thing works, and it ended with a six-week-old outage nobody had reported, a feature brought back from the dead, and a lesson about which signal to believe.

I · The smallest question

A change had shipped. The tests were green, the pull request was merged, the deploy went out, and everyone moved on. The only thing left was a question so ordinary it felt like a formality: how do I actually see this working. Not read the diff. Not trust the checkmark. See the thing do the thing, in production, with my own eyes.

That question is the whole story. Everything that follows was already sitting there, in production, waiting. All it took to surface was one person refusing to accept a green checkmark as proof that a thing was alive.

II · Green is not the same as working

The feature I wanted to validate needed a real run to prove itself, the kind that only happens when a signed-in person actually uses the product end to end. So I asked for exactly that, and the run failed. Not with a clean error, but with the worst kind, the intermittent kind: a slow screen, a timeout, a generic message that told me nothing about why.

A merge being green means the code compiles and the tests the code carries with it pass. It does not mean the running system, with its real database and its real network and its real production configuration, does what the code describes. Those are two different claims, and the gap between them is exactly where things go to hide. The checkmark was honest about what it measured. It measured the code. Nobody had measured the system.

A green merge is a statement about the code. Whether the system is alive is a separate question, and it only gets answered by someone insisting on the real run.
III · Two gaps that hid each other

Pulling on the failed run led straight down. Every signed-in attempt to save a decision was returning a server error, and so was every attempt to load an old one. Reads and writes, both broken. This was not a hiccup. This was the product's core loop, down, in production.

Then the second gap, the one that explains the silence. The last decision the system had successfully saved was dated six weeks earlier. Not because nobody had tried since, but because saving had been broken that entire time and no one had said a word. The reason no one said a word is its own quiet finding: almost nobody finishes a run all the way to a saved outcome. Two out of two hundred forty-seven. A loop that few people complete is a loop that can break in total silence, because a break produces no complaints when there is barely any traffic to complain.

The RCA, floor by floor

Symptomsigned-in saves and reads both return 500 in production; new users see a timeout, not an error. First readevery failure is the same database error: a column the code expects does not exist in the live database. Why silentthe last saved decision is six weeks old; the save loop is completed so rarely that a total outage drew zero reports. Root causea schema migration shipped in the code but never ran against production. The code and the database had quietly disagreed for six weeks.

Two gaps, each invisible on its own, each the reason the other stayed hidden. The outage hid because the loop is rarely finished. The rarely-finished loop hid because nothing forced the outage into view.

IV · The column that was never there

The root cause was almost dull, which is the point. A migration, written weeks ago, added a column to a table. It was committed, reviewed, merged, deployed. Every checkmark it needed to pass, it passed. But the production database had never been placed under the tool that applies migrations on deploy. So the migration lived in the code, describing a column that did not exist in the database it described, and nothing in the pipeline ever noticed the two had drifted apart.

This is the trap in its pure form. Code and database are two artifacts that must agree, and nothing was checking that they did. A migration file is a promise. A promise that is never executed is indistinguishable, from the code's point of view, from one that was. The only way to catch it is to ask the database itself what it actually contains, and until that failed run, nobody had asked.

V · Trust the artifact, not the signal

Here is the turn that made the day worth writing down. Once the column was restored, I went to confirm the fix by having that same slow run tried again. And again the screen said it had failed. Load failed, the same generic message as before. My first instinct was that the fix had not worked.

The instinct was wrong, and being wrong about it taught me more than the outage did. I went to the database instead of the screen, and there it was: a freshly saved decision, timestamped to the minute of the run the screen had just called a failure. The work had completed. The record was real. The client had simply given up waiting before the server finished, and reported a failure the server never had.

The saved record is the truth. The screen is a hint. A client that calls a slow success a failure is describing its own patience, not the state of the world.

So the failure the user sees and the failure that actually occurred are two different things, and I had nearly trusted the first as the second. The screen is downstream of a network that can drop, a browser that can time out, a phone on a weak signal. The saved artifact is downstream of none of that. When they disagree, the artifact wins, every time.

VI · What came back with it

The one-line fix that restored the missing column did something I did not expect. It brought a whole feature back from the dead. There was an annotation layer, built weeks earlier, fully wired, with real logic and real interface, that had never once worked in production. It depended on the exact table the missing migration was supposed to create. So every time anyone had tried to use it, it had failed in the same silent way, for the same reason.

The moment the migration ran, the feature simply started working. Buttons that would have thrown an error an hour earlier now did what they were built to do, and the records they wrote landed in the table that finally existed. A feature is not shipped when its code merges. It is shipped when a real person can use it against the real system. By that measure this one had never shipped at all, and reviving it cost nothing but the same fix the outage already needed.

VII · So the next one is loud

A bug found by luck is a bug that will be found by luck again, later, worse. The only honest response to a failure that hid for six weeks is to make its whole class impossible to hide next time. Not fix the instance. Fix the silence.

So the fix was not only the missing column. It was a diagnostic the system did not have: one request that reports, in one place, whether the database is reachable, whether the code and the schema still agree, how long since the last saved record, and whether the dependent services answer. The twenty minutes of log reading it took to find this becomes one call. And the timeout that cried failure over a real success now recovers instead: when the screen loses patience, it goes and looks for the saved record before it dares tell you it failed.

The controls that came out of it

Diagnostica health endpoint that probes the live schema against what the code expects, plus reachability, record freshness, and dependent services. The buried RCA becomes one request. Recoveryon a client timeout, the app polls for the record that may have saved anyway and routes to it, instead of reporting a failure the server never had. Preventionmigrations applied on deploy, so the code and the database can no longer drift apart in silence. The queued durable fix. Receiptsthe outage fix, the diagnostic, and the recovery each logged against the real record: restored column, real saved rows, verified live.

Make the class loud, not the instance. A break that survives six weeks did not lack a fix. It lacked a way to be seen.

VIII · The question was the tool

None of this was found by a monitor, an alert, or a test. It was found by a person asking the smallest honest question, how do I know this works, and then refusing every answer that was not the real system doing the real thing. The checkmark said yes. The code said yes. The database, when finally asked, said no, and had been saying no for six weeks to anyone who bothered to ask, which was no one.

That is the whole of it. Validation is not the ceremony you perform after the work is done. It is the one move that finds what every green signal is structured to miss, because a green signal reports on the thing it watches, and the failures that matter live precisely in the gap between the things being watched. Ask the smallest question. Then go and look at the thing itself.

One honest footnote on method. The walk from symptom to cause here was not improvised: reproduce the real error, separate symptom from cause, establish pre-existing versus introduced, fix minimally, verify by driving the failure path, capture the root cause so it cannot hide again. ARCS ships that exact discipline as an open skill. This incident was run by hand, but step for step it is that method, which is the entire reason a method is worth writing down: so the next person under pressure does not have to invent it.

The method, as an open skill

rca-skilla Claude Code skill that turns "something broke" into disciplined root-cause analysis: triage the effort, reproduce the real error, fix the cause not the symptom, verify by driving the failure path, and capture the root cause. Try itfeatured on good.arcs.care; source at github.com/flashesofbrilliance/rca-skill.

Run by hand this time. The point of the skill is that it never has to be.

More field notes Read the rest of the series →