Teaching Everyone to Fish for Tokens 教会每个人为Token而渔

Housekeeping: No voiceover for this post as I’m traveling. 说明:由于我正在旅行,本文章没有配音。

The oldest comparison people try to make is how what’s happening with open models compares to foundational open-source software projects like the Linux operating system. There are fairly clean analogies, but they paint a narrow path forwards for the self-sustaining nature of the open-source model ecosystem, where once Linux got big enough it was going to be self-fulfilling as the best possible tool for many jobs. The open-source language model – i.e. only models that come with a full training recipe, data, code, etc. – is a closer analogue to the open-source operating system. The open weight models you use – those with just model weights and inference code to run them – are closer to specific versions of software that you install in a project built upon them. 人们最常做的比较是,当前开放模型(open models)的发展与Linux操作系统等基础开源软件项目有何相似之处。有一些相当清晰的类比,但它们在开源模型生态系统的自我维持性方面描绘了一条狭窄的前进道路——一旦Linux足够庞大,它就会自我实现,成为许多任务的最佳工具。开源语言模型——即只提供完整训练配方、数据、代码等的模型——更接近开源操作系统。而你所使用的开放权重模型(open weight models)——那些只包含模型权重和运行它们的推理代码——更接近于你在基于它们构建的项目中安装的特定软件版本。

Share 分享

Model weights are very transient on average, but they still have a long shelf life, as with a lot of heavily used software. It’s why many companies are still using workflows built on Llama 3, despite agentic behaviors taking off years later. The open-source recipe, typified in modern times by the Olmo models I helped build at Ai2, with its predecessors like Pythia from EleutherAI, is a resource intensive process that any company can pick up, modify, and press “run” on to produce a new set of model weights. In the best cases, the community can contribute improvements in data or training code back into the next model too! This is why Nvidia is investing so much in nearly open-source models – for their Nemotron models they release all the data they legally can and the training code, etc. Nvidia wants a world where countless people can build token machines, so intelligence is not monopolized. This is a world with massive demand for inference across many companies, all of which want to buy Nvidia’s offerings. 模型权重平均而言非常短暂,但它们仍有很长的保质期,就像许多广泛使用的软件一样。这就是为什么尽管智能体行为在几年后才兴起,许多公司仍在使用基于Llama 3构建的工作流。开源配方,现代以我在Ai2帮助构建的Olmo模型为代表,其前身如EleutherAI的Pythia,是一个资源密集型过程,任何公司都可以拾起、修改,然后按“运行”按钮,生成一组新的模型权重。在最好的情况下,社区还可以将数据或训练代码的改进贡献给下一个模型!这就是为什么Nvidia在近乎开源的模型上投入如此之多——对于他们的Nemotron模型,他们发布了所有法律允许的数据和训练代码等。Nvidia希望创造一个无数人都能构建“token机器”的世界,从而让智能不被垄断。这是一个在许多公司中对推理有巨大需求的世界,所有这些公司都想购买Nvidia的产品。

Open-source AI has a tricky future, as building the best models is extremely capital intensive. The ability to build competitive models has stayed more accessible in industry longer than many would’ve expected. The default expectation for many is that training models is too expensive and the open-source recipe is too far behind, so building a new lab centered on some part of training LLMs will not be tractable. 开源AI的未来充满变数,因为构建最好的模型需要极其密集的资本。在业界,构建有竞争力模型的能力比许多人预期的要保持得更可及。许多人的默认预期是,训练模型过于昂贵,开源配方远远落后,因此建立一个专注于训练LLM某一部分的新实验室将是不可行的。

There are two futures from here. First is if “it works” – if the open-source recipe works for Nvidia, they’ll be creating far more demand for their chips (and profits) than it costs to build the models. Right now it’s

reported

that Nvidia is spending $26 billion on this endeavor. It’s not clear if this will work, or if AI’s capital intensiveness will drive more and more companies out of the training game. We

haven’t seen many signs of this starting

. In fact, the companies bowing out – like

Databricks

and 01.ai – seem like anomalies. 从这里开始有两条未来路径。第一条是“它奏效了”——如果开源配方对Nvidia奏效,他们将为芯片创造远超构建模型成本的需求(和利润)。目前据报道,Nvidia正为此投入260亿美元。尚不清楚这是否会奏效,或者AI的资本密集性是否会将越来越多的公司挤出训练游戏。我们还没有看到很多这种开始的迹象。事实上,像Databricks和01.ai这样退出的公司似乎是反常现象。

The open-source ecosystem will become increasingly dependent on Nvidia’s financing in the coming years. This is an existential window, where within a few years the profits of this approach need to return to them, or another open model company needs to cultivate platform-like financial feedback loops on their openness. This economic reward needs to be proportional to the profits generated by Anthropic and OpenAI’s APIs to keep pace over decades of language model development. This can be driven by competitiveness on performance or by the AI boom just being so big that the open model training, inference, and fine-tuning companies all have vast quantities of demand. 在未来几年,开源生态系统将越来越依赖Nvidia的资助。这是一个关乎存亡的窗口期,在几年内,这种方式的利润需要回报给他们,或者另一家开放模型公司需要在其开放性上培育类似平台的财务反馈循环。这种经济回报需要与Anthropic和OpenAI的API所产生的利润成比例,才能在数十年的语言模型发展中跟上步伐。这可以由性能上的竞争力驱动,或者由AI繁荣规模之大,使得开放模型的训练、推理和微调公司都拥有巨大需求来驱动。

The second future is if one of these two financially positive paths doesn’t play out, open models will fork to a different development path than the leading closed models – one more focused on efficiency, modifiability, specialization, etc. I put this mentally as

my most likely outcome

– open models are still incredibly useful, but

fill a long-tail ecosystem

relative to the closed counterparts that have monopoly ownership stakes in the most valuable areas like knowledge work collaboration, drug discovery, SWE, etc. The long-tail is something like enterprise-specific agents that run on-prem with private data on repetitive business tasks. 第二条未来路径是,如果这两条财务上积极的路径都未能实现,开放模型将分叉到一条与领先闭源模型不同的发展路径——一条更注重效率、可修改性、专业化等的路径。我在心里将这一点视为最可能的结果——开放模型仍然非常有用,但相对于在知识工作协作、药物发现、软件工程等最有价值领域拥有垄断性所有权利益的闭源对手,它们填补了一个长尾生态系统。长尾指的是类似企业专用智能体的东西,它们在本地运行,使用私有数据处理重复性业务任务。

Part of why I think this open-source training will have a hard time catching on is because training is getting more complex and more abstracted. The current open model ecosystem is buoyed by an explosion in interest in post-training open models. These people take models like DeepSeek V4 Flash, Inkling Small, or GLM 5.X and finetune them for their specific agentic tasks (e.g. in Tinker, the most popular finetuning API today). 我认为开源训练将难以流行起来的部分原因是,训练正变得越来越复杂和抽象。当前的开放模型生态系统得益于对后训练开放模型兴趣的爆发。这些人采用DeepSeek V4 Flash、Inkling Small或GLM 5.X等模型,并针对他们的特定智能体任务进行微调(例如在Tinker中,这是目前最流行的微调API)。

Interconnects AI is a reader-supported publication. Consider becoming a subscriber. Interconnects AI是一份由读者支持的出版物。请考虑成为订阅者。

Over the last few years, post-training largely referred to the whole process of modifying the base model to make it intelligent and usable. There is a shift happening where the ability to train a base model to be a general agentic reasoner is becoming opaque like at-scale pretraining practices from a few years ago. This could go so far as to change the established pretraining, midtraining, post-training lexicon that has been standard for a few years. It could come to be something closer to pretraining, reasoning training, and post-training. 在过去几年中,后训练很大程度上指的是修改基础模型使其智能且可用的整个过程。现在正在发生一种转变,即训练基础模型成为通用智能体推理器的能力正变得像几年前的规模化预训练实践一样不透明。这甚至可能改变多年来一直标准的既定预训练、中训练、后训练词汇。它可能会变得更接近预训练、推理训练和后训练。

As there’s less interest in training the entire model, there’s less interest in investing in open-source AI. These are the only sort of hints we will get, but we cannot do much to fight the economic gravity of these situations. This trend is the next step in the number of open model builders who release base models (the model versions before core reasoning training) continuing to decrease. It goes hand in hand with open model builders experimenting with

revenue

share

licenses for downstream use in products or inference. These are experiments in keeping the financing viable for building near frontier open-weight models – a lot hinges in the near future on how successful they are. These are the people that need to succeed for Nvidia’s demand-growth strategy around open-source to succeed, and last. 随着对训练整个模型的兴趣减少,对投资开源AI的兴趣也在减少。这是我们能得到的所有暗示,但我们无法对抗这些情况的经济引力。这一趋势是发布基础模型(核心推理训练之前的模型版本)的开放模型构建者数量持续减少的下一步。它伴随着开放模型构建者尝试在产品或推理的下游使用中采用收入分成许可。这些是为了保持构建接近前沿的开放权重模型融资可行的实验——在不久的将来,很多取决于这些实验的成功程度。这些是需要成功的人,才能让Nvidia围绕开源的需求增长战略成功并持久。

Along the way we’re still in for a ton of action in open-weight models, as releasing access to intelligence is one of the strongest business strategies available. This additional type of player, who monetizes the AI indirectly, is typified by Meta and other hyperscalers with massive balance sheets. Meta

releasing its very-strong Muse Spark 1.2 model

as open-weights would severely hamper the revenue growth rate of their competitors in Anthropic and OpenAI who rely on selling tokens. These companies are both commoditizing their complements, but they’re doing it in different ways. Nvidia wants to teach everyone to fish for tokens, so the ecosystem is self-sustaining, but Meta is strategically flooding the zone with tokens. 在这一过程中,我们仍将看到开放权重模型的大量动作,因为释放智能访问权是最强大的商业策略之一。这种额外的玩家类型,通过间接方式将AI变现,以Meta和其他拥有巨额资产负债表的超大规模企业为代表。Meta发布其非常强大的Muse Spark 1.2模型作为开放权重,将严重削弱其竞争对手Anthropic和OpenAI的收入增长率,后者依赖出售token。这些公司都在将自己的互补品商品化,但方式不同。Nvidia想要教会每个人钓鱼(token),使生态系统自我维持,而Meta则在战略性地用token淹没整个区域。