Today I Learned · Calibration Receipts
Three predictions · three falsifiers · one correction · Jan–May 2026

Closing the Loop

Three dated, public calls about Apple, each with a stated falsifier attached in advance. Here is how they resolved, including the one number I had wrong.
score before → check after → correct in public
Why keep the receipts

Predictions are only worth the falsifier attached to them. Anyone can be right about Apple in hindsight; the harder discipline is writing down, before the fact, exactly what would prove you wrong, then checking back once the fact is no longer avoidable. Between January and May 2026 I published three pieces arguing Apple's apparent weaknesses were being misread. Each named a specific, checkable condition under which it would fail. This is the accounting.

Three claims, three falsifiers, checked against the record

3 predictions scored before the fact · 1 correction owned · 0 retrofitted

None of the language below was written after the fact; the falsifier text is quoted from what was published at the time. Every claim resolved with a caveat attached, and one of those caveats is a correction, not a footnote.

Prediction 1 · CEO succession · dated Jan 30, 2026
Claim
Ternus named Cook's presumptive heir, reading the Q.ai deal as a hardware-mandate signal.
Falsifier
A directional call, not numeric: no named successor by a stated date.
Resolution
Transition announced Apr 20, 2026; Ternus became CEO effective Sep 1; Cook moved to executive chairman.
Status
CONFIRMED
Prediction 2 · gross-margin floor · dated Feb 19, 2026
Claim
Unified-memory architecture holds gross margin above the stated floor through the memory-cost supercycle.
Falsifier
"If Apple's gross margins fall materially below 47% in Q2 2026 due to memory costs, the argument weakens significantly."
Resolution
Q2 FY26 gross margin: 49.3%. Q3 FY26: 50.1%, including roughly a 2-point tariff-refund tailwind.
Status
CONFIRMED, with a correction owned below
Prediction 3 · platform control, not cosmetic · dated May 25, 2026
Claim
iOS 27 "Extensions" opens real third-party model routing, not a cosmetic compatibility layer.
Falsifier
Extensions ships as "a walled choice menu," not a real developer ecosystem.
Resolution
Confirmed at WWDC 2026 (June 8): a system-wide framework routing Apple Intelligence, Siri, Writing Tools and Image Playground to Gemini, Claude and ChatGPT.
Status
CONFIRMED as real, not graded on reach

The one number I had wrong

An earlier internal draft read the Q2 gross margin as 48.7%. That is wrong. The reported figure, confirmed against MacRumors' coverage of the April 2026 call and Apple's own newsroom release, is 49.3%. I am correcting it here rather than quietly fixing it and moving on, because the whole point of a calibration exercise is that the correction is part of the record, not an embarrassment to route around.

The Q3 figure carries its own caveat, more important than the Q2 typo. The 50.1% reported for Q3 FY26 included an estimated 2-point favorable impact from tariff refunds, a one-off. Strip that and the underlying number is closer to the low-to-mid 48% range: still above the 47% floor, but by a narrower margin than the headline implies. The falsifier held. The size of the win was smaller than the unadjusted number suggests, and saying so is what makes the confirmation worth trusting.

What "confirmed" does not mean

Two distinctions, both toward restraint. The succession call was directional, made roughly seven months before the effective date, naming a person rather than a mechanism or a date. It resolved as stated; it should not be read as more precise than it was.

The Extensions call is confirmed as a real, system-wide routing framework announced at WWDC. It is not confirmed as general availability to every user, and this piece does not claim that. What resolved is the architecture question, not the adoption curve.

Pre-mortem: the Convex Skeptic

The obvious objection is that this is survivorship dressed up as calibration: three hits, no visible losses, prove only that the author remembers the calls that landed. That deserves a direct answer.

Skeptic
Three wins out of an unknown private denominator is cherry-picking dressed as rigor.
Record
All three were public, dated and falsifiable at publication, not selected afterward. The falsifier language is verbatim from the originals.
Skeptic
A generous author picks a falsifier easy to clear.
Record
The 47% margin line and the "walled choice menu" outcome were both real, specific things Apple could have shipped instead. Neither was a coin flip, and the Q2 correction shows the check was run against primary sources, not asserted.
Skeptic
Three points over four months on one company is too thin to call a "system."
Record
Correct, and this does not call it a system. It is a decision architecture: dated priors, stated falsifiers, and a public correction when the record and the draft disagreed. The method is what is demonstrated, not a track record.

What this does not show

What this does not prove is whether stated conviction tracks how often the call is right. Three resolved bets is not a calibration curve; it is three points. The same self-graded confidence rubric runs on these pieces already: the "AI Strategy Wasn't Late" article capped its own falsifiability at 12 out of 20, because two of its four falsifiers had not resolved at publication. That discipline closes three predictions here. It does not yet prove the rubric predicts accuracy at any reliable rate, which takes a real sample size across a real range of outcomes, losses included. And to be exact about the standard: this is a single-observer, self-graded record, not a peer-reviewed or independently-replicated efficacy study, and I do not claim it is. The honest reading of this page is a demonstrated method, not a track record.

The difference between a hot take and a decision architecture is not confidence. It is whether the falsifier was written down before anyone knew the answer.