【文章标题】:OpenAI Jalapeño: Better Than Nvidia Blackwell 【文章标题】:OpenAI Jalapeño:优于英伟达 Blackwell
【文章正文】: 【文章正文】:
OpenAI has spent the past couple years quietly building “Jalapeño,” an inference chip just announced at Hot Chips. Rumors of a successful tapeout had been swirling for a while. But now we have details. OpenAI invited us to look at their chip,
go to their labs to check out how real it is, and
benchmark
it with our
InferenceX
suite.
OpenAI 在过去几年里一直在悄悄打造“Jalapeño”,这是一款刚刚在 Hot Chips 上发布的推理芯片。关于成功流片的传言已经流传了一段时间。但现在我们有了详细信息。OpenAI 邀请我们查看他们的芯片,去他们的实验室验证其真实性,并使用我们的 InferenceX 套件进行基准测试。
In June,
OpenAI unveiled the chip program
in partnership with Broadcom, built from a blank slate exclusively for LLM inference.
6 月,OpenAI 与博通合作公布了芯片计划,从一张白纸开始专为 LLM 推理打造。
Design work began in the middle of 2024
, going from initial team hiring to manufacturing tape-out in ~16 months, an extremely fast ASIC development cycle.
设计工作始于 2024 年年中,从最初团队招聘到制造流片仅用了约 16 个月,这是一个极快的 ASIC 开发周期。
In general first generation chips are not competitive, but OpenAI bucks the trend by being industry leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models. OpenAI does this with extreme hardware software codesign. Surprisingly, OpenAI is not over specialization on any specific part of model inference, but instead by focusing on being a general chip that delivers high performance in all scenarios.
一般来说,第一代芯片并不具备竞争力,但 OpenAI 打破了这一趋势,成为行业领先者,并在多个顶级开源模型上击败了我们能够测试的所有英伟达、AMD 和谷歌芯片。OpenAI 通过极致的软硬件协同设计实现了这一点。令人惊讶的是,OpenAI 并没有过度专注于模型推理的某个特定部分,而是专注于成为一款在所有场景下都能提供高性能的通用芯片。
In this article, we will go into architectural details, software details and performance results for Jalapeño on InferenceX.
在本文中,我们将深入探讨 Jalapeño 在 InferenceX 上的架构细节、软件细节和性能结果。
Source:
OpenAI
来源:OpenAI
A generalized inference chip
一款通用推理芯片
Everyone says that OpenAI’s chip is specialized for OpenAI models, but that’s wrong, OpenAI made a generalized chip for AI inference.
每个人都说 OpenAI 的芯片是专为 OpenAI 模型定制的,但这是错误的,OpenAI 制造了一款用于 AI 推理的通用芯片。
The timelines are insane. It shows that claims that use of AI is being used to accelerate chip design are real. Regardless of the quick timelines,Open AI spent a bunch of money, made pragmatic design decisions and their team is cracked, so this comes as no surprise.
时间线令人难以置信。这表明使用 AI 加速芯片设计的说法是真实的。尽管时间线很短,OpenAI 投入了大量资金,做出了务实的设计决策,而且他们的团队非常出色,因此这并不令人意外。
Just looking at the specs, it is an immediate contender:
仅看规格,它立即成为一个有力的竞争者:
Source: SemiAnalysis
来源:SemiAnalysis
And the use of HBM4 makes it stand out as comparable to flagship GPUs from NVIDIA and AMD:
而 HBM4 的使用使其脱颖而出,可与 NVIDIA 和 AMD 的旗舰 GPU 相媲美:
Source: OpenAI
来源:OpenAI
A lot of the media coverage of this chip has followed a few throwaway comments from OpenAI that claim the chip will be optimized for their models in a way that other chips are not. This is wrong. Jalapeño is a generalized inference chip capable of running all sorts of models, and all sorts of workloads, including our benchmark InferenceX, where we ran the benchmark with OpenAI engineers in the lab. As a joke, OpenAI even showed us it running Doom, which was ported to their chip with just Codex prompts.
许多媒体对这款芯片的报道都沿用了 OpenAI 的一些随口评论,声称该芯片将针对其模型进行其他芯片所没有的优化。这是错误的。Jalapeño 是一款通用推理芯片,能够运行各种模型和各种工作负载,包括我们的基准测试 InferenceX,我们在实验室中与 OpenAI 工程师一起运行了该基准测试。作为玩笑,OpenAI 甚至向我们展示了它在运行《毁灭战士》(Doom),而这款游戏仅通过 Codex 提示就被移植到了他们的芯片上。
The following is our headline perf/W result, looking at token throughput per All-in utility MW.
以下是我们最重要的每瓦性能结果,考察的是每兆瓦总设施功耗下的令牌吞吐量。
Jalapeño smokes every other chip
. All this is done without Multi Token Prediction (MTP), while the other chips on the chart are the best performing configs of each respective SKU, all with MTP.
Jalapeño 完胜其他所有芯片
。这一切都是在没有多令牌预测(MTP)的情况下完成的,而图表中的其他芯片是各自 SKU 中性能最佳的配置,且都使用了 MTP。
Source: SemiAnalysis
来源:SemiAnalysis
Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point in the curve. It excels not only in low-latency scenarios but also in high-throughput scenarios. A more apples to apples comparison is against Single Token Prediction results, it knocks every competitor out of the water. At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model.
Jalapeño 在几乎所有场景下的每瓦性能都击败了 Blackwell,而且并未针对曲线上的任何特定点进行调优。它不仅在低延迟场景中表现出色,在高吞吐量场景中同样优异。更公平的比较是与单令牌预测结果对比,它让所有竞争对手都相形见绌。在低并发场景下,Jalapeño 展现出卓越的交互性,在 DeepSeek R1 模型上并发数为 1 时,每用户每秒可超过 700 个令牌。
Incredibly, this is all achieved with single-token prediction (STP), no speculative decoding and no prefill-decode disaggregation. In addition to DeepSeek R1, we also got to see some other models, including Kimi-K2.5 and GPT-OSS which ran at approximately 1,400 tok/sec/user. For all models, we confirmed that
Jalapeño’s
GSM8k evals attained results on par with Nvidia chips.
令人难以置信的是,这一切都是在单令牌预测(STP)下实现的,没有推测解码,也没有预填充-解码分离。除了 DeepSeek R1,我们还看到了其他一些模型,包括 Kimi-K2.5 和 GPT-OSS,它们的运行速度约为每用户每秒 1,400 个令牌。对于所有模型,我们确认 Jalapeño 的 GSM8k 评估结果与英伟达芯片相当。
Some caveats on this. First, all numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of
InferenceX
benchmarks nor have we seen
AgentX
results. AgentX is our preferred suite for comparing chip performance due to the datasets’ long context and multi-turn characteristics that reflect the cache behavior of realistic production workflows. Frameworks that perform well on 8k1k may perform worse on AgentX as real production loads stress components like routers, prefix cache mechanisms, cache management, offload infrastructure, etc. These are not tested by single turn 8k1k. Read more about this in out AgentX article.
对此有一些注意事项。首先,所有数据均由 OpenAI 提供给我们。我们亲自在实验室验证了 InferenceX 的运行,但我们没有运行完整的 InferenceX 基准测试套件,也没有看到 AgentX 的结果。AgentX 是我们比较芯片性能的首选套件,因为其数据集具有长上下文和多轮次特征,能反映现实生产工作负载的缓存行为。在 8k1k 上表现良好的框架在 AgentX 上可能表现更差,因为真实生产负载会对路由器、前缀缓存机制、缓存管理、卸载基础设施等组件造成压力。这些都不是单轮 8k1k 能测试到的。请在我们的 AgentX 文章中了解更多信息。
Second, we believe that comparison to Blackwell is somewhat incomplete and unfair. Jalapeño is really competing against chips like Rubin that also use HBM4. Vera Rubin systems are starting to ship to customers right now, while it will still be some time before OpenAI has anything beyond engineering samples of Jalapeño.
其次,我们认为与 Blackwell 的比较有些不完整且不公平。Jalapeño 实际上是在与同样使用 HBM4 的 Rubin 等芯片竞争。Vera Rubin 系统目前已经开始向客户发货,而 OpenAI 要拿出 Jalapeño 工程样品之外的产品还需要一段时间。
Thus, performance should really be compared against Rubin, not Blackwell, and in some sense we expect a custom chip like Jalapeño to outperform Blackwell. Vera Rubin NVL72 delivers 5.4x the perf/MW of GB200 NVL72
as we described in our article analyzing the NVIDIA performance claims in their launch with CoreWeave last month
. We will compare Jalapeño to Vera Rubin’s Ju
因此,性能实际上应该与 Rubin 进行比较,而不是 Blackwell,从某种意义上说,我们预计像 Jalapeño 这样的定制芯片会超越 Blackwell。Vera Rubin NVL72 的每瓦性能是 GB200 NVL72 的 5.4 倍,正如我们在上个月分析 NVIDIA 与 CoreWeave 联合发布中的性能声明文章中所描述的那样。我们将把 Jalapeño 与 Vera Rubin 的 Ju