Why catching up is cheaper than leading in AI
Nathan Lambert's recap of the open model wave makes a point that is easy to miss under the model-release noise: catching up is structurally cheaper than leading. He quotes an engineer explaining the edge plainly, "it helps because we're not trying to push the frontier, we are just trying to catch up." A team that knows a capability is possible, and roughly how it was built, can skip the expensive dead ends that the first lab had to explore. That, more than raw compute, is why the gap between the best open and closed models has stayed within a few months.
He also pushes back on a popular explanation for how the fast followers do it. Ben Thompson argued that distillation, training a smaller model on a stronger one's outputs, matters more during the reinforcement learning phase. Lambert disagrees on economics and evidence. Using a frontier model like Claude or GPT to grade 20 to 40 million RL rollouts would be far too expensive to run at scale, and he notes the research literature does not cleanly show that a stronger teacher yields a better student after supervised fine-tuning. He calls it an open question rather than a settled trick. On the near term he expects no models above 10 trillion parameters this year, with the open frontier held by Kimi, GLM 5.2, DeepSeek, and Qwen.
Why it matters
If you plan around open models, the takeaway is that the few-month gap is likely to hold, so betting your stack on open weights is not the risk it looked like a year ago. And if you were counting on distillation as the reason cheap models keep closing in, Lambert's argument is a reason to check that assumption before you build a strategy on it.