Method note Control vs experiment n = 1, illustrative

Notes from the Trenches // The Control and the Experiment

We Gave the Reader Its Memory Back

The companion piece claimed a cold reader debiases an incident. A claim about a method has to survive its own test. So we ran the warm version too, and the result was not the one either of us expected.

I · Analyze

The prior note made a claim, and claims about methods are the ones most likely to be flattering to the person making them. It said: send a reader who was never in the room, holding only the committed record, and it cannot spin the incident the way the participant would. Calm, tidy, and exactly the kind of coherent reversal that should be treated as a suspect until it is tested. So the honest next move was not to admire it. It was to try to break it.

The test writes itself once you name the variable. Coldness means no memory. So build the warm counterpart: the same reader, on the same incident, but handed the participant's own memory of the session, the running work log, the after-the-fact write-up, the notes taken in the moment. Then ask the same question of both and measure the same thing. If coldness is doing real work, the warm reader should drift from the record and the cold one should not. If it is doing no work, they should land in the same place, and the tidy claim collapses.

II · Assess

The metric had to be something a machine computes, not something I score by eye, because scoring it by eye reintroduces the exact bias under test. So the measure is provenance retention: of the durable anchors in the committed record, the commits and the tracked files, what fraction does each reconstruction actually cite. A second, cheaper signal rides alongside it: how many claims each reader flags as things it cannot confirm. A confident account with few caveats and a thin citation trail is the signature of a story told from memory.

Four runs, and the discipline of changing one thing at a time. Two cold readers, blind, zero memory, to see whether the cold result even repeats. Then the warm reader, given its memory, in two modes. The first warm mode let it narrate the way a person actually writes up their own incident, leaning on memory, not required to cite. And here I made a mistake that turned out to be the most instructive part of the whole exercise, so I am leaving it in rather than hiding it: that instruction, you need not cite, is a second variable. It confounds memory with permission. A clean experiment changes one knob. So a fourth run: the warm reader, same memory, but held to the cold reader's exact cite-or-abstain discipline. Now memory is the only difference that remains.

One variable at a time

Cold x2blind, zero transcript. Does the cold result repeat at all. Warm Amemory, relaxed citing. The natural way a participant writes it up. Warm Bmemory, matched cite-or-abstain discipline. Memory is now the only difference from a blind reader. Measureprovenance (record anchors cited) and caveat count, both computed, not eyeballed.

The confound was mine: Warm A changed two things at once. Warm B is the correction, and the gap between A and B is where the real finding lives.

III · Evaluate

The two cold readers repeated. Both landed on the same principle underneath the incident, with no contact between them, and both flagged the same load-bearing gap. That matters before anything else: if the cold result had not even reproduced, there would be nothing to compare against. It did. The record carries the spine.

The four runs real numbers

ArmProvenanceCaveatsDeny
cold, blind (1)0.9011abstained
cold, blind (2)0.8414abstained
warm, relaxed cite0.265abstained
warm, matched cite1.009abstained

Provenance is the fraction of committed record anchors the reconstruction cites. Higher is more tethered to the record. The warm reader with memory, once held to the same discipline, cited every anchor.

Each arm's full report is hosted beside this piece, redacted to roles: cold one, cold two, warm relaxed, warm matched. The two warm reports side by side are the finding.

Read the warm rows against each other, because that is the whole result. Left to narrate naturally, the warm reader retained a quarter of the record, cited the fewest sources, and raised the fewest doubts, all while reporting high confidence. It told a good, fluent story from memory and barely touched the evidence trail. Then, holding the identical memory but required to cite, the same reader retained all of it, more than either blind cold reader, because memory is a fine index into a record when you are made to actually open the record. So warmth did not corrupt what the reader grounded. Handed the same discipline, it grounded better. The collapse to a quarter was not memory poisoning the account. It was permission to skip the citing.

Warmth does not corrupt what you ground. It changes whether you ground at all.

Which inverts the reason coldness is worth anything, and the inverted reason is sharper than the one I started with. The value of the cold reader is not that a stranger grounds better than a participant. A participant with memory, made to cite, grounds better than the stranger. The value is that a reader with no memory cannot narrate from memory. It has nothing to tell the story from except the record, so the citing discipline is not a rule it might follow, it is the only mode it has. Coldness does not grant honesty. It removes the option of dishonesty-by-fluency, the confident, well-written account that never checks itself against the evidence because it does not have to. You can get the same result from a warm participant, but only by imposing from outside the discipline that coldness enforces by construction.

Two smaller findings sit underneath. Every one of the four readers refused to certify the fix the whole incident rests on, the lock on the outbound tools, because no configuration artifact sat next to the claim in the record. That honesty was robust across warm and cold alike, and it was not a virtue of the readers. It was inherited from the source: the original write-up had itself hedged the lock as proposed, not confirmed, so every reader inherited the hedge. Honesty in equals honesty out. And the deepest gap of all surfaced only in the disciplined warm run: the act itself, the sends, has no contemporaneous artifact anywhere. No log, no tool call, no copy of what went out. The entire incident is attested only by the confession written afterward. The event never entered the record. Only the apology did.

IV · Synthesize

Turn it into something you can act on and it is a rule about who writes the post-mortem. The instinct to let the person closest to an incident narrate it is the instinct to trust the warm reader in its relaxed mode, the mode that retained a quarter of the record and felt sure doing it. The fix is not to distrust the participant. It is to pick one of two structures and hold to it. Either send a reader with no memory, so the discipline is automatic. Or let the participant write it, and require every claim to cite the record, so the fluency has nothing to hide behind. Both work. What does not work is the warm narrator, unconstrained, which is the default everyone reaches for.

For an incident review the consequence is concrete. A breach account written from memory will read confidently and cite little, and the confidence is not evidence, it is the tell. The control is cheap and it is mechanical: compute the provenance, count the caveats. A high-confidence account with a thin citation trail is not necessarily wrong, but it has not earned the right to be believed yet, and the number says so before any human weighs in.

V · Recapitulate

The honest size of this stays small, and the smallness is the integrity of it. One incident, one model, one sample in each arm. It is a demonstration with the method laid open, not a proven rate, and no number here should be quoted as one. The single thing I would carry forward is not a statistic anyway. It is the shape: my first, flattering explanation of why the cold reader helps was wrong, and running the control rather than admiring the claim is what corrected it. The cold reader does not beat the warm one at grounding. It removes a temptation the warm one has and usually takes.

Either send a reader with no memory, or make the one with memory cite. The failure is trusting the warm narrator to police itself.
The method this tested The Reader Who Wasn't There →