The Open Source AI Singularity Was Supposed to Be Here by Christmas. It’s Not.
I almost retweeted a chart at 2am last week. Open-source LLMs closing the gap with closed models — benchmarked on the Artificial Analysis Intelligence Index, trend line pointing straight at December 2026. Six months to parity. Open-source AI utopia by Christmas.
Looked solid. Almost hit send.
Then I pulled up the other seventeen benchmarks.
Yeah. Ouch.
Eighteen total. The gap between open and closed sits at five months. Hasn’t moved in two years. Not getting wider either, just… stuck.
Except — and this is worth saying — coding.
Coding is the one place that actually changed. Fifteen months behind to maybe one or two. If you’ve been waiting to run a real coding model locally, no API bills eating your budget, no round-trip latency killing your flow — this is actually your moment. No asterisk required.
Everything else? Reasoning, factual stuff, instruction following, creative writing — same gap as two years ago. Maybe a little worse in some cases, I honestly can’t tell without spending real time in the eval methodology and who has that kind of time. The picture isn’t flattering either way.
Here’s the part nobody in the open-source camp wants to hear: pick your benchmark, pick your story. Intelligence Index says open source is about to win. The other seventeen say open source has been consistently behind with almost no progress. And — not shocking — the benchmarks that circulate on Twitter are the ones that look good. Every LLM company cherry-picks. It’s not malicious, it’s just how press releases work.
So. The honest version.
Open source is genuinely closing in coding. That’s real, measurable, and replicable — I ran some of these benchmarks myself on a 4090 and the numbers held up. Everywhere else you’re looking at five, sometimes six months behind frontier closed models and that gap hasn’t meaningfully changed in two years.
Better than where we were. But the viral chart told a different story — and now you know why it spread so fast.
What the Numbers Actually Say
The Artificial Analysis Intelligence Index has been making the open-source case since mid-2024. Their trend line projects parity around winter 2026. Makes for a great tweet. Terrible summary of what’s actually happening.
Across eighteen benchmarks — MMLU, HumanEval, MATH, GSM8K, LiveCodeBench, the whole gang — the average gap has sat between 4.5 and 5.5 months for about two years. Flat. Not compressing. Not diverging either, just… there.
Until you look closer.
Coding benchmarks tell a completely different story. HumanEval, MBPP, LiveCodeBench — open models went from 12-to-15 months behind to 1-to-3 months. That’s real. It shows up every time, across different evaluation setups, and you can run it yourself in an afternoon if you’ve got a halfway decent GPU.
The rest of the suite? Flat. Slightly worse in some cases. Nobody’s putting those charts on Twitter.
Why You’re Only Seeing the Good Charts
Think about it from a marketing perspective. Eighteen benchmarks. Three look great, fifteen are Meh or Worse. Which ones are you putting in your blog post?
Closed-source companies do exactly the same thing, for what it’s worth. This isn’t an open-source-specific problem. The Intelligence Index happens to look better for open models than most alternatives because — wait for it — that’s literally what they built the index to track. Kind of circular when you think about it.
I’m not saying the numbers are made up. I’m saying they’re curated. There’s a difference.
The benchmarks that circulate are the ones that landed well. The rest live in appendices, buried in arXiv papers nobody reads, or just not mentioned. This is how the whole industry communicates — not just AI, honestly. Software has always worked this way. The changelog is always better than the product.
The One Place It Actually Worked
Coding is the outlier and I think it’s because coding benchmarks are almost impossible to fudge. You run HumanEval, you get a number. Twenty minutes, one GPU, reproducible. No human raters arguing about whether the output “feels” right. The test is the test.
Compare that to instruction following or factual accuracy — benchmarks where annotation quality, evaluator bias, and test set contamination introduce enormous variance. You could spend six months debating whether your model got better at following complex instructions or whether your eval set just got easier. Nobody has the budget for that debate, so what survives is what has clean measurement. Coding wins because the scoreboard is unambiguous.
Open source moved fast in coding because the game was fair. Everywhere else the rules keep changing and the scoreboard’s blurry.
The Thing Nobody Wants to Say
Open source LLMs are a way better deal today than two years ago. For coding. If you’re self-hosting. The economics, the latency, the control — all improved meaningfully.
But the “open source is catching closed AI” story is mostly the “coding is catching closed AI” story wearing a better headline. The reasoning gap, the knowledge gap, the general capability gap — those haven’t moved.
I almost retweeted that chart. I’m supposed to be the person who checks these things before sharing. If I almost fell for it, the Product Manager who saw the headline and shared it to the team Slack definitely did. Worth remembering the next time a clean chart shows up in your feed with a misleading thread attached.
Anyway. Need more coffee.