Mozilla puts the open-vs-closed gap at 4.4 months — and warns the premium buys longer jobs, not better ones
Mozilla's State of Open Source AI puts Chinese open-weights models 4.4 months behind the US frontier and asks whether the closed model is worth paying for.

Mozilla published the second edition of its State of Open Source AI report on September 15, 2026, showing it to Ars Technica before release. The exclusive ran the next day under a headline that folded the argument into one line: paying for frontier AI models buys a four-month head start at five times the cost. The report's own framing is narrower: the capability distance between the best Chinese open-weights models and the US frontier has closed to about 4.4 months, and Mozilla's practical conclusion is that most organizations should use open models as the default for the majority of their work.
The 4.4 months is a composite reading, not a stopwatch
It is not a measurement of anything a user would recognize. The number is read off composite index scores, and the index underneath it is Artificial Analysis's Intelligence Index, where Moonshot AI's Kimi K3 lands within three points of Anthropic's closed Fable 5 while costing roughly 30 percent as much. Three index points and four months are different units; the conversion is Mozilla's arithmetic, unpublished in a form an outsider can rerun. Composite indexes also move when constituents are reweighted or a new evaluation is folded in, so a gap expressed in months inherits the instability beneath it — and that index was revised twice inside a single week in September.
Two trackers, two numbers, one moving target
Epoch AI's data insight on the US-China gap puts Chinese models an average of seven months behind the US frontier over the period it covers, with a minimum gap of four months and a maximum of 14. That does not contradict a 4.4-month reading taken at one moment; it shows how much the answer depends on when you look and what you count.
| Measurement | Reading | What it actually covers |
|---|---|---|
| Mozilla composite index gap | 4.4 months | Best Chinese open weights vs US frontier closed models, one moment |
| Closed-model long-horizon advantage | 1.7 times, ~12 hours vs ~7 | Best model of each kind at a 50 percent reliability threshold |
| Epoch AI average US-China lag | 7 months, range 4 to 14 | Chinese models vs the US frontier over the period covered |
The premium buys longer jobs
The report's most concrete claim is not about average quality. Mozilla CTO Raffi Krikorian's illustration to Ars uses the time-horizon measure, the length of expert task a model completes at a 50 percent reliability threshold: "If the open frontier can handle a seven-hour job, the closed frontier can handle a 12-hour one. In four months, the open model handles the 12-hour job, and the closed one handles something around 20." On that measure the best closed model currently manages a job 1.7 times as long as the best open one. The gap is concentrated in eight-to-12-hour tasks, nothing in the comparison reliably handles more than 12 hours, and sub-eight-hour work can go to the cheaper open model. "We see the decision to pay for closed models as workload-specific rather than organization-specific," Krikorian said, adding that a closed model "earns its premium in a few places: expert professional work, high-intensity retrieval, and long context." And on when to pay: "Pay when that head start is worth it... Routine work you'll still be doing next quarter is not, because you'll be able to do it for a fifth of the cost soon, and the model won't be the bottleneck anyway."
The cost gap comes from somebody else's harness
The five-times figure does not come from Mozilla's index. Benchmarking company Vals AI ran models on its own harness in the Terminal-Bench 2.1 evaluation, where Z.ai's open-weights GLM 5.2 scored within one point of Anthropic's Claude Opus 4.7 and 4.8 while costing about five times less per completed task. Ars flags the caveat: closed labs often ship harnesses tuned for their own models, which can flatter them on someone else's. One neutral harness is the strongest class of evidence in this argument, and still one data point. Z.ai's line of open releases, including the 5.3 generation, keeps resetting the comparison.
Popular is not the same as profitable
Adoption and revenue point in opposite directions. The report notes that eight of the top 10 models on OpenRouter by token volume in August 2026 provide open weights. A Linux Foundation paper by Frank Nagle and Daniel Yue found open models earned 4 percent of revenue against 96 percent for closed models, on data from May to September 2025. Ars notes that DoorDash uses Kimi for routine work while reserving Fable for harder tasks, matching the split Krikorian describes.
Whose capital keeps the lane open
Krikorian is blunt that the plural ecosystem is more concentrated than it looks. "The current reality is that most of the open models the world runs on are Chinese," he said, and "The Chinese labs are running the same playbook the Americans ran with Android—give it away, but own the ecosystem around it." The funding picture is "the uncomfortable truth": "That's a plurality and a concentration at the same time." His prescription has US and European labs competing in "the same open lane, so that no single country sets the world's defaults", backed by public compute for fully open reference models, foundations holding neutral ground, companies that benefit from commodity models paying in, and philanthropy covering evaluation and audit. On transparency: "It's hard to fully trust a model with decisions if you can't tell how it was trained or what it was evaluated against."
What open weights do not include
Mozilla's numbers describe open weights, not open source. Anyone can download the main components, but training data, the training pipeline and the training code are typically withheld — so the 4.4-month gap is about what is downloadable, not what is reproducible.
What would settle it
Everything above is directional. The composite-index gap is one reading at one moment; the time-horizon ratio is a different measurement in different units; the five-times cost figure is a third party's run on one benchmark. Repeated open-versus-closed runs on the same neutral harness at matched budgets, with per-task cost curves published next to the scores instead of a single average. An independent replication of Mozilla's month-gap arithmetic, including the index weighting that turns three points into four months. Until then the defensible claim is narrower: for work shorter than a shift and not on a deadline, the price gap is real and the quality gap may not be worth paying for.


