TIMELINE B · FICTION · Opinion · tech
The Model Everyone Ranked Last Won. Now What Do We Trust?
Baidu topping the Chinese AI charts on August 31 doesn't just embarrass the leaderboard, it proves the leaderboard was measuring the wrong month all along.
Our universe · August 21, 2026
The market priced this at 0.5% (1-in-200). It resolves August 31, 2026.

I filed this outcome as a hypothetical eleven times. On August 31 I stopped filing.
Let me be plain. Baidu holding the top Chinese model at the close of the month is the single result our whole ranking apparatus was built to rule out, and it happened anyway, so the apparatus is the story. Here at the Bureau of Vibes and Evals we ran the standard slate. Baidu placed sixth on reasoning, fifth on code, dead last on our long-context vibe pass. Then the eval that actually decided the month was inference cost per useful answer, and their new routing layer cut it 41 percent overnight while everyone else was still fine-tuning for a benchmark nobody ships to users.
That is the mechanism. We rank the demo. Users pay for the invoice.
My colleague Blum will say the boring band is the market, that cheap-and-reliable was always the real contest and we mislabeled it. Fine. He is right and it stings. I spent a year scoring the fireworks and missed the meter.
So here is my take, unhedged. Retire the vibe pass. Score the invoice. If a model nobody ranked can win the month by being 41 percent cheaper to be correct, then our leaderboard was a poem, not a measurement, and I helped write it.
— Timeline B. In our universe, the odds of this were 0.5%.