Predictions are only worth the falsifier attached to them. Anyone can be right about Apple in hindsight; the harder discipline is writing down, before the fact, exactly what would prove you wrong, then checking back once the fact is no longer avoidable. Between January and May 2026 I published three pieces arguing Apple's apparent weaknesses were being misread. Each named a specific, checkable condition under which it would fail. This is the accounting.
None of the language below was written after the fact; the falsifier text is quoted from what was published at the time. Every claim resolved with a caveat attached, and one of those caveats is a correction, not a footnote.
An earlier internal draft read the Q2 gross margin as 48.7%. That is wrong. The reported figure, confirmed against MacRumors' coverage of the April 2026 call and Apple's own newsroom release, is 49.3%. I am correcting it here rather than quietly fixing it and moving on, because the whole point of a calibration exercise is that the correction is part of the record, not an embarrassment to route around.
The Q3 figure carries its own caveat, more important than the Q2 typo. The 50.1% reported for Q3 FY26 included an estimated 2-point favorable impact from tariff refunds, a one-off. Strip that and the underlying number is closer to the low-to-mid 48% range: still above the 47% floor, but by a narrower margin than the headline implies. The falsifier held. The size of the win was smaller than the unadjusted number suggests, and saying so is what makes the confirmation worth trusting.
Two distinctions, both toward restraint. The succession call was directional, made roughly seven months before the effective date, naming a person rather than a mechanism or a date. It resolved as stated; it should not be read as more precise than it was.
The Extensions call is confirmed as a real, system-wide routing framework announced at WWDC. It is not confirmed as general availability to every user, and this piece does not claim that. What resolved is the architecture question, not the adoption curve.
The obvious objection is that this is survivorship dressed up as calibration: three hits, no visible losses, prove only that the author remembers the calls that landed. That deserves a direct answer.
What this does not prove is whether stated conviction tracks how often the call is right. Three resolved bets is not a calibration curve; it is three points. The same self-graded confidence rubric runs on these pieces already: the "AI Strategy Wasn't Late" article capped its own falsifiability at 12 out of 20, because two of its four falsifiers had not resolved at publication. That discipline closes three predictions here. It does not yet prove the rubric predicts accuracy at any reliable rate, which takes a real sample size across a real range of outcomes, losses included. And to be exact about the standard: this is a single-observer, self-graded record, not a peer-reviewed or independently-replicated efficacy study, and I do not claim it is. The honest reading of this page is a demonstrated method, not a track record.
The difference between a hot take and a decision architecture is not confidence. It is whether the falsifier was written down before anyone knew the answer.