Development
Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
September 15, 2026 Development Source: Ars Technica
Share this article
But the steadily narrowing performance gap between open and closed frontier models, coupled with the cost-effectiveness of open models, is now widely recognized by many companies. For example, delivery company DoorDash has been using Kimi for routine work while reserving Fable for more difficult tasks that would normally take human experts longer to complete.
Unsurprisingly, this has accelerated the use of open models since Mozilla’s inaugural State of Open Source AI report was published on July 14.
The narrowing performance gap is being measured in several ways. For example, the research nonprofit METR has defined an AI model’s time horizon as the length of tasks—measured by how long human experts require to complete them—that can be handled by AI models with a “reliable” 50 percent success rate. That time horizon has been doubling on a set cadence that has accelerated over time.
The best closed model can currently do a job that is 1.7 times as long as the longest job that the best open model can reliably finish.
“If the open frontier can handle a seven-hour job, the closed frontier can handle a 12-hour one,” Krikorian told Ars. “In four months, the open model handles the 12-hour job, and the closed one handles something around 20.”
Tasks requiring between eight and 12 hours are typically the ones that a closed frontier model can do and the open models cannot handle yet, Krikorian said. No models are generally capable of reliably doing tasks longer than 12 hours, whereas tasks taking less than eight hours can be done by either type of model and could be handed off to the cheaper open models.
There are many nuances and caveats for such comparisons. For example, closed frontier models often come with their own harness—the software layer that helps the models access various tools and memory to perform agent-like actions. Harnesses custom-built by AI labs for their models can sometimes boost performance on various tasks, whereas the models may perform less well with third-party harnesses.
The current reality is that “most of the open models the world runs on are Chinese,” Krikorian said. “The Chinese labs are running the same playbook the Americans ran with Android—give it away, but own the ecosystem around it,” he explained.
But Krikorian is not so concerned about “Chinese models” as he is about the risk of concentration. At the moment, the best open models are concentrated in China, whereas the best closed frontier models come from US companies. He argued for US and European labs to start competing in the “same open lane, so that no single country sets the world’s defaults.”
“The uncomfortable truth is that the plural ecosystem around open-weight AI is largely funded by Chinese capital right now,” Krikorian said. “That’s a plurality and a concentration at the same time.”
The history of development for open source software like Linux may offer some lessons for how an “alternative coalition” could come together, Krikorian said. Such open source software infrastructure was funded by a coalition of organizations that “each needed the commodity layer to exist,” including neutral foundations.
Any coalition for building more open AI models would likely consist of “institutions with a mission rather than a market,” as opposed to frontier AI labs, Krikorian said. He described it as follows:
We need public compute programs funding fully open reference models. Switzerland’s national compute producing Apertus is an example. Foundations need to hold the same neutral ground that they held for the internet’s open protocols, extended into models, harnesses, and agentic standards. Companies that benefit from commodity models have to participate. And philanthropy has to cover what none of the others will—evaluation and audit infrastructure, especially.
Krikorian also called for the alternative coalition to go even further than open-weights models by embracing more of the open source approach. “It’s hard to fully trust a model with decisions if you can’t tell how it was trained or what it was evaluated against,” he said.