TIMELINE B · FICTION · Opinion · tech
I Told You the Benchmark Was Rigged in Favor of the Underdog. Nobody Listened.
The month OpenAI reclaimed the top slot proves my long thesis: leaderboards reward whoever ships the boring reliability nobody wanted to fund.
Our universe · August 5, 2026
The market priced this at 1.8% (1-in-54). It resolves August 31, 2026.

I have spent three years arguing that model supremacy is decided by plumbing, not genius. On August 31 the Bureau of Vibes and Evals crowned OpenAI's model the single best on Earth, and my inbox filled with people who owe me a coffee.
Here is my one argument. The lead did not come from a smarter model. It came from a 41-second drop in median latency after they rebuilt their scheduling stack around Halcyon Compute's spare capacity. Speed moved the vibe score. Nothing else did.
"We re-ran the eval twelve times," said Dr. Ines Okafor-Halvorsen, our chief vibes officer. "The model that answers before you finish sighing wins the human-preference tiebreak. Every time."
Everyone wanted the story to be a breakthrough. It was a warehouse in Ohio with better cooling.
The lesson the industry will ignore, again, is that finishing beats founding. The Institute of Applied Finishing has said this about strikers for a decade. It applies to matmul too.
So spare me the think-pieces about emergent reasoning. The best model this month is the one that stopped keeping people waiting. I called it. Buy me the coffee.
— Timeline B. In our universe, the odds of this were 1.8%.