Notes from the Trenches // A calibration receipt
Three dated, public calls about Apple, made between January and May 2026, each with a stated falsifier attached in advance. Here is how they resolved, including the one number I had wrong.
score before → check after → correct in public
Predictions are only worth the falsifier attached to them. Anyone can be right about Apple in hindsight; the harder discipline is writing down, before the fact, exactly what would prove you wrong, then checking back once the fact is no longer avoidable. Between January and May 2026 I published three pieces arguing that Apple's apparent weaknesses (a "behind on AI" narrative, a memory-cost headwind, a succession question) were being misread by the coverage. Each piece named a specific, checkable condition under which the argument would fail. This is the accounting.
The first, published January 30, 2026 ("$2B Poker Face"), read Apple's Q.ai acquisition as a hardware-mandate signal and named John Ternus, then head of Hardware Engineering, as Tim Cook's "presumptive heir." The second, published February 19, 2026 ("The Bloomberg Op-Ed Is Wrong About Tim Cook"), argued that Apple's unified-memory architecture would let gross margin hold above a stated 47% floor through the memory-cost supercycle, and set that 47% line as the explicit falsifier. The third, published May 25, 2026 ("Apple's AI Strategy Wasn't Late. The Coverage Was."), argued that Apple's iOS 27 "Extensions" system would open real third-party model routing rather than functioning as, in the piece's own words, "a walled choice menu", and pointed at the coming WWDC keynote as the test.
All three are linked in the footer. None of the language below was written after the fact; it is quoted from what was published at the time.
Prediction 1 · CEO succession · dated Jan 30, 2026
Prediction 2 · gross margin floor · dated Feb 19, 2026
Prediction 3 · platform control, not cosmetic · dated May 25, 2026
Three for three against falsifiers stated in advance, before the outcome was known. None resolved without a caveat attached, and one of those caveats is a correction, not a footnote.
The Q2 gross-margin figure in an earlier internal draft of this piece read 48.7%. That number is wrong. The reported figure, confirmed against MacRumors' coverage of the April 2026 earnings call and Apple's own Q3 newsroom release, is 49.3%. I am correcting it here rather than quietly fixing it and moving on, because the entire point of a calibration exercise is that the correction is part of the record, not an embarrassment to route around.
The Q3 figure carries its own caveat, and it is a more important one than the Q2 typo. The 50.1% gross margin reported for Q3 FY26 included an estimated 2-percentage-point favorable impact from tariff refunds, a one-off. Presented without that disclosure, 50.1% reads as pure structural margin strength. It is not. Strip the tariff refund and the underlying number is closer to the low-to-mid 48% range, still above the stated 47% floor, but by a narrower margin than the headline figure implies. The falsifier held. The size of the win was smaller than the unadjusted number suggests, and saying so is what makes the confirmation worth trusting.
Two more distinctions matter, both in the direction of restraint rather than overclaim.
The succession call was directional, made roughly seven months before the effective date, naming a person rather than predicting a mechanism or a date. It resolved as stated. It should not be read as more precise than it was: the original piece named a likely heir on the strength of an internal capital-allocation signal, not a leaked transition memo.
The Extensions call is confirmed as a real, system-wide routing framework, announced at WWDC. It is not confirmed as general availability to every user, and this piece does not claim that. The falsifier named in the original article was whether Extensions would be "a real developer ecosystem" or "a walled choice menu." What resolved is the architecture question, not the adoption curve. Anyone citing this piece for a claim about current end-user rollout is citing it for something it does not say.
The obvious objection to a piece like this is that it is survivorship dressed up as calibration. Three hits, presented after the fact, with no visible losses, prove nothing except that the author is good at remembering the calls that landed and quiet about the ones that did not. That objection deserves a direct answer, not a wave.
The skeptic's strongest version of the objection is not answered by these three predictions alone. It is answered by whether the same discipline, falsifier stated before the fact, correction owned after it, holds on the next three.
The difference between a hot take and a decision architecture is not confidence. It is whether the falsifier was written down before anyone knew the answer.
This is not a claim that the underlying method produces certainty. Two of the three claims (the succession call, the Extensions architecture call) were directional reads of publicly visible incentive structures, not statistical forecasts, and the sample here is three predictions across four months on one company. What it demonstrates is narrower and, I think, more useful in a principal-investor conversation than a scorecard: that a falsifier stated in advance and a correction issued in public are cheaper to produce than most people assume, and that doing both changes what a prediction is worth to the person reading it afterward.
What it does not show is whether stated conviction tracks how often the call is right. Three resolved bets is not a calibration curve; it is three points. The same self-graded confidence rubric runs on these pieces already: the "AI Strategy Wasn't Late" article capped its own falsifiability at 12 out of 20, because two of its four falsifiers had not resolved at publication. That discipline closes three predictions here. It does not yet prove the rubric predicts accuracy at any reliable rate, which takes a real sample size across a real range of outcomes, losses included. The honest reading of this page is a demonstrated method, not a track record.
The next test of the same discipline is already running. Whatever gets published next carries the same rule: name the falsifier before the outcome is known, and if a number in the draft turns out wrong, say so in the same place the wrong number appeared.
The discipline that makes the series legible is the same whether the subject is a custody clause or a piece of slang: name the mechanism, then state plainly what would prove you wrong.
Each entry states its own falsifier. Read together they are the topography: not one call, but the inference quality across the series.