【文章标题】:Nvidia’s AI advantage is moving beyond the GPU
【文章标题】:英伟达的AI优势正在超越GPU
【文章正文】:
Before this week, the dominant story about Nvidia went something like this: For the first few years of the AI boom, Nvidia was the only source for state-of-the-art GPUs, which became immensely profitable as the industry scaled out. In the last few years, hyperscalers like Amazon and Google have started building their own chips, and Nvidia is no longer the only game in town, leading many investors to wonder how durable its advantage really is.
本周之前,关于英伟达的主流叙事是这样的:在AI热潮的头几年,英伟达是尖端GPU的唯一供应商,随着行业规模扩大,这些GPU变得极其赚钱。过去几年,亚马逊和谷歌等超大规模企业开始自研芯片,英伟达不再是唯一玩家,这让许多投资者质疑其优势的持久性。
It’s a compelling story, and mostly true. After growing its market cap 10x between the start of 2023 and mid-2025, Nvidia shares have been on a more modest trajectory for the past year, driven by concerns about GPU competition.
这个说法很有说服力,且基本属实。在2023年初至2025年中市值增长10倍后,过去一年英伟达股价因GPU竞争担忧而走势趋缓。
A new narrative has taken shape since the company’s earnings on Wednesday and investors are starting to realize that Nvidia’s advantage goes far beyond GPUs. As AI’s compute grows into the gigawatt scale, orchestration has become an increasingly complex task. Not surprisingly, Nvidia has built much of the state-of-the-art hardware needed to handle it, giving the company a huge advantage in the systems that surround the GPU even as it sees increased competition on the GPUs themselves.
周三财报发布后,新的叙事逐渐成形:投资者开始意识到英伟达的优势远不止GPU。随着AI算力迈入吉瓦级,系统协调已成为日益复杂的任务。不出所料,英伟达已构建了处理该任务所需的大量尖端硬件,使其在GPU外围系统领域获得巨大优势——尽管GPU本身正面临更激烈竞争。
For all the talk of compute as a commodity, it’s still incredibly difficult to operate a megascale data center at peak efficiency — and as deployments get bigger and faster, that challenge is only growing.
尽管人们常将算力视作大宗商品,但让超大规模数据中心保持峰值效率仍极其困难——随着部署规模扩大、速度提升,这一挑战只会加剧。
Rack by Rack
逐机柜突破
You can see some of this just by looking at the details of what Nvidia is actually selling. The company is currently rolling out its Vera Rubin architecture, which pairs the Rubin GPU with a collection of other units, including the Vera CPU, the Groq 3 LPX inference accelerator and similar racks for storage and networking.
从英伟达实际销售的产品细节即可窥见一斑。该公司正在推出Vera Rubin架构,将Rubin GPU与Vera CPU、Groq 3 LPX推理加速器及存储网络专用机柜等组件整合。
Over the past week, I’ve been talking to folks at Nvidia about what those systems actually do, and the results have been surprising. Like the Rubin GPU itself, they’re extremely specialized systems, but instead of churning through tokens, they’re making sure everything outside the GPU works as efficiently as possible. If the GPU is the engine, these are the rest of the car.
过去一周我与英伟达员工探讨这些系统的实际功能,结果令人惊讶。它们和Rubin GPU一样高度专业化,但其职责并非处理数据,而是确保GPU外围组件以最高效方式运作。如果说GPU是引擎,这些系统就是整辆车的其他部件。
The Vera CPU in particular is focused on the problem of orchestrating data. “Vera is important because there’s only so much memory that you can put in a single server or any sort of compute platform,” Jason Hardy, Nvidia’s VP of storage technology, told me.
Vera CPU尤其专注于数据协调问题。英伟达存储技术副总裁Jason Hardy告诉我:“Vera之所以重要,是因为单台服务器或任何计算平台的内存容量都存在物理上限。”
As data centers have scaled up computing power, memory capacity has scaled up too, which is why companies like Micron have gotten rich in the second wave of the infrastructure boom. But getting that data to the GPU at the right time isn’t straightforward — and as companies look to drive tokens-per-watt lower and lower, they’re realizing how important that kind of traffic direction is.
随着数据中心算力提升,内存容量同步增长,这正是美光等企业在基础设施第二波热潮中获利的原因。但如何适时将数据送达GPU并非易事——当企业追求每瓦特处理效能持续优化时,它们越发认识到这种数据调度的重要性。
“We saw upwards of 3x improvement in these operations, where the Vera CPU is allowing for acceleration,” Hardy said. “So now we can use our flash to its fullest potential, because we can get all that performance out of it without bottlenecking.”
Hardy表示:“Vera CPU使这些操作的性能提升达3倍以上。现在我们能充分发挥闪存潜力,因为可以在无瓶颈情况下榨取全部性能。”
You can see versions of the same problem outside of Nvidia. When OpenAI developed its Jalapeño chip, a major focus was avoiding these challenges entirely by minimizing the amount of data that needs to be moved around.
类似问题在英伟达之外同样存在。OpenAI开发Jalapeño芯片时,核心思路就是通过最小化数据迁移量来彻底规避这些挑战。
“We designed Jalapeño to minimize data movement and communication delays,” the company said in a blog post earlier this month. “Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end.”
该公司本月早些时候在博客中表示:“Jalapeño的设计宗旨是减少数据迁移和通信延迟。其大算域使整个工作负载保持在互联系统内,最小化数据移动,确保请求全程快速高效。”
It’s a different approach, avoiding data movement entirely by conducting a workload within one integrated chip. But the overall logic is the same, increasing efficiency with smarter traffic control instead of just more processor cycles. That in turn opens up a whole new layer of infrastructure for companies to compete over.
这是种不同路径——通过单芯片集成处理彻底避免数据迁移。但底层逻辑一致:用更智能的流量控制(而非更多处理器周期)提升效率。这为企业在基础设施领域开辟了全新竞争维度。
This new focus on data orchestration isn’t automatically a win for Nvidia. The company will have to compete with rival chipmakers and hyperscalers just as it has with GPUs. But the competition has moved to a new layer, where building a rival GPU matters less than being able to make the entire system work efficiently.
数据协调的新焦点不会自动成为英伟达的胜利。该公司将像应对GPU竞争一样,与芯片制造商和超大规模企业展开较量。但竞争已升级至新层面——打造竞品GPU的重要性,已逊色于让整个系统高效运作的能力。
And at least in the early stages, Nvidia looks to have a commanding lead.
至少在现阶段,英伟达似乎拥有压倒性优势。