【文章标题】:Are Open Models Catching Up?

【文章标题】:开源模型正在迎头赶上吗?

【文章正文】:

【文章正文】:

The past two months have been a breakout period for open source AI. Yes, there was the “DeepSeek moment” back in January 2025, but no one actually used R1 to do any economically valuable work. In contrast, models like

GLM 5.3 and Kimi K3 are genuinely capable of many of the same coding and agentic tasks that rocketed Anthropic to $65B+ ARR

. Unlike others who inflated ARR,

our figures were much closer to reality.

Source: SemiAnalysis

过去两个月是开源 AI 的爆发期。是的,2025 年 1 月曾有过“DeepSeek 时刻”,但当时没有人真正使用 R1 来完成任何具有经济价值的工作。相比之下,像

GLM 5.3 和 Kimi K3 这样的模型真正能够胜任许多相同的编码和智能体任务,正是这些任务让 Anthropic 的 ARR 飙升至 650 亿美元以上

。与其他夸大 ARR 的公司不同,

我们的数据要接近现实得多。

来源:SemiAnalysis

It is an exciting time to be a token consumer. Competition is heating up, usage resets are being doled out, and the battle for your tokens now extends beyond the OpenAI-Anthropic duopoly. Fireworks alone is processing over

40T

tokens per day—2x the OpenAI API’s

volume

at the end of March.

作为 token 消费者,这是一个激动人心的时刻。竞争正在升温,使用量重置正在发放,争夺你的 token 的战场如今已不再局限于 OpenAI-Anthropic 双头垄断。仅 Fireworks 一家每天处理的 token 数量就超过

40T

——是 3 月底 OpenAI API

用量

的两倍。

However,

major FUD has also emerged as a result of open model success

: if open models stay capable enough relative to the closed frontier at a fraction of the cost, won’t the model layer become commoditized? This outcome would obviously be disastrous for frontier lab margins. For full details on Anthropic and OpenAI’s financials, see our

Tokenomics Model

.

然而,

开源模型的成功也引发了严重的 FUD

:如果开源模型能够以极低的成本保持相对于闭源前沿模型的足够能力,模型层难道不会商品化吗?这一结果显然会对前沿实验室的利润率造成灾难性影响。有关 Anthropic 和 OpenAI 财务的完整详情,请参阅我们的

Tokenomics 模型

。

To project how the open vs closed capability gap will progress in the future, we first need to measure the past. Naively, you might pick a single set of benchmarks to measure all historical models, but this is a mistake.

要预测开源与闭源能力差距在未来将如何演变,我们首先需要衡量过去。天真地看,你可能会选择一组基准来测量所有历史模型,但这是一个错误。

Every benchmark is a product of a particular era.

每个基准都是特定时代的产物。

When someone creates a new benchmark, their goal is to discern differences in model capabilities at the time. If they’re successful, the model makers will climb said benchmark until it becomes saturated. Once that happens, everyone stops caring about the benchmark, and the cycle repeats.

当有人创建一个新基准时,他们的目标是辨别当时模型能力的差异。如果他们成功了,模型制造者就会不断攀登该基准,直到它饱和。一旦发生这种情况,所有人都会不再关心这个基准,循环再次重复。

There have been three eras thus far in the history of LLMs: early scaling, reasoning, and agentic

. Each era represented a step-function increase in model utility, and rather than trying to plot a single continuous trend, we believe it’s better to evaluate the models and benchmarks from each era individually.

迄今为止,LLM 的历史上已经历了三个时代:早期扩展、推理和智能体

。每个时代都代表了模型实用性的阶跃式提升,我们认为与其试图绘制一条连续的总体趋势,不如分别评估每个时代的模型和基准。

When viewed this way, it becomes clear that

the open vs. closed gap moves in cycles

. At the start of each era, a frontier lab completes some promising research, trains an impressive model, deploys it at scale to their users, and jumps ahead. Then, other labs identify the key advances, reverse-engineer what the frontier lab is doing, replicate them in their own models, and close the gap. Nothing stays secret forever—especially when you factor in distillation. It’s just a question of how long it takes.

从这个角度看,很明显

开源与闭源的差距呈周期性变化

。在每个时代之初,某个前沿实验室完成了一些有前景的研究,训练出一个令人印象深刻的模型,将其大规模部署给用户,并取得领先。随后,其他实验室识别出关键进展,逆向工程前沿实验室的做法,在自己的模型中复制这些进展,并缩小差距。没有什么能永远保密——尤其是当你把蒸馏考虑在内时。这只是需要多长时间的问题。

To answer this question, we took all the relevant models from each era and ran a curated set of benchmarks to get a composite capability score.

为了回答这个问题,我们选取了每个时代所有相关模型,并运行了一组精心挑选的基准,以获得综合能力得分。

The result is a clear trend: with each generation, open-source models take half as long to catch up to the first closed-source model of the era.

结果呈现出明显的趋势:每一代开源模型追上该时代第一个闭源模型所需的时间都会缩短一半。

Source: SemiAnalysis

来源:SemiAnalysis

Of course, benchmarks don’t tell the full story, and we’ll highlight all the relevant caveats below. Finally, we’ll extend this analysis into the future, and explain why it’s less bearish frontier models than you might initially think.

当然,基准并不能说明全部情况,我们将在下文强调所有相关的注意事项。最后,我们会将这一分析延伸到未来,并解释为什么它对前沿模型的看空程度可能比你最初想象的要低。

How we measured

我们如何衡量

Here’s an overview of the models and benchmarks we selected for each era:

以下是我们为每个时代选择的模型和基准的概览:

Source: SemiAnalysis

来源:SemiAnalysis

Picking a single SOTA closed and open model at a particular time is subjective, but our selections reflect the general consensus among AI experts. In cases where there’s debate—e.g. Fable 5 vs GPT 5.6 today—we were conservative and tested both.

在特定时间挑选单一的 SOTA 闭源和开源模型是主观的,但我们的选择反映了 AI 专家的普遍共识。在存在争议的情况下——例如今天的 Fable 5 与 GPT 5.6——我们采取了保守做法,对两者都进行了测试。

For benchmarks, we relied on a combination of personal taste and popularity. Humanities Last Exam (HLE), for example, is known to have lots of

issues

, but was also truly one of the defining benchmarks of the reasoning era with no close substitutes. SWE-bench Pro, on the other hand, is similarly popular and

problematic

, but also closely approximated by DeepSWE.

对于基准,我们结合了个人偏好和流行度。例如,Humanities Last Exam(HLE)虽然已知存在许多

问题

,但它确实是推理时代最具定义性的基准之一,且没有相近的替代品。另一方面,SWE-bench Pro 同样流行且

存在问题

,但 DeepSWE 可以很好地近似它。

Most of the benchmark scores here we ran ourselves using

Prime Intellect

‘s evaluation stack, specifically their environments hub and the evals harness included in

Prime-RL

. The rest come from runs by our friends at

Artificial Analysis

and Datacurve’s

DeepSWE leaderboard

. Open models were served the way they would have been at release: vLLM versions, hardware that was in use at the time, and sampling settings from the model card. For closed models, we ran against their pinned API versions. Where our numbers share a chart with third-party values, we matched their rulesets.

这里的大多数基准分数是我们使用

Prime Intellect

的评估栈自行运行的,具体包括他们的环境中心和

Prime-RL

中包含的评估框架。其余分数来自我们在

Artificial Analysis

的朋友以及 Datacurve 的

DeepSWE 排行榜

。开源模型按照发布时的方式提供服务:vLLM 版本、当时使用的硬件以及模型卡中的采样设置。对于闭源模型,我们针对其固定的 API 版本运行。当我们的数据与第三方数值出现在同一图表中时,我们匹配了他们的规则集。

We’d like to give a huge thank you to Florian Brand (

@xeophon

) from Prime Intellect

for helping us pick benchmarks/models, implement evals, and check for correctness.

我们要衷心感谢 Prime Intellect 的 Florian Brand(

@xeophon

),感谢他帮助我们选择基准/模型、实施评估并检查正确性。

Era 1 | Early scaling (2022-2024)

时代 1 | 早期扩展(2022-2024)

It’s June 2023. The world is reckoning with ChatGPT, and Mark Zuckerberg just agreed to fight Elon Musk at the Colosseum. But while Zuck is training jiu-jitsu and doing Murphs, his company is doing some training of their own. FAIR is about to push past the Mistral exodus and other drama, and successfully ship Llama-2-70B. The first open model that approached the frontier.

那是 2023 年 6 月。世界正在适应 ChatGPT,马克·扎克伯格刚刚同意在斗兽场与埃隆·马斯克对决。但就在扎克伯格训练柔术和做 Murph 训练时,他的公司也在进行一些自己的训练。FAIR 即将挺过 Mistral 团队出走和其他风波,成功发布 Llama-2-70B。这是第一个接近前沿水平的开源模型。

How far behi

落后多