【文章标题】:Cerebras’s Next Generation CS-4: Fast Just Got Faster 【文章标题】:Cerebras新一代CS-4:快上加快

【文章正文】: Cerebras revealed CS-4 this week, with more details to come at Hot Chips. CS-4 is their fourth-generation rack built around the same third generation 5nm wafer-scale engine: WSE-3. CS-4 doubles the performance of CS-3 through increased power consumption and clock frequency per wafer, and better rack-scale density. 【文章正文】: Cerebras本周发布了CS-4,更多细节将在Hot Chips上公布。CS-4是围绕第三代5nm晶圆级引擎WSE-3构建的第四代机架。通过提高每晶圆的功耗和时钟频率,以及更好的机架级密度,CS-4的性能是CS-3的两倍。

This all translates into CS-4 being able to double the tokens/s/user per wafer from CS-3, and at around the same cost as the previous generation. This is a no-brainer for customers who can enjoy double the token revenue with the same hardware spend. It’s not only the tokens that are getting faster, but so is time to market: the rack architecture itself is redesigned to be more modular, allowing improved manufacturability and deployment times. Last but not least, a new I/O module will enable open, heterogeneous and disaggregated inference architectures going forward. These disaggregated inference setups will go a long way to help overcome the memory capacity constraints of the CS-4 by pairing it with HBM-based systems. 这一切意味着,CS-4每晶圆能够提供的每秒每用户token数(tokens/s/user)是CS-3的两倍,而成本与上一代大致相同。对于客户来说,这是一个不假思索的选择——同样的硬件支出,却能获得双倍的token收入。不仅仅是token变快了,上市时间也缩短了:机架架构本身被重新设计为更加模块化,从而提升了可制造性和部署速度。最后但同样重要的是,一个新的I/O模块将使未来的开放、异构和解耦推理架构成为可能。这些解耦推理设置通过与基于HBM的系统配对,将大大有助于克服CS-4的内存容量限制。

Source: SemiAnalysis 来源:SemiAnalysis

Same Wafer, Double the Clock 同一晶圆,双倍时钟

Source: SemiAnalysis 来源:SemiAnalysis

The CS-4 uses the same 5nm WSE-3 as the CS-3, but Cerebras is extracting double the performance by doubling clock speeds. This comes from feeding dramatically more power to the wafer, and is enabled by the CS-4’s improvement in power delivery and cooling technology. While staying on the same 5nm silicon sounds underwhelming, Cerebras can still double the metric that matters most: memory bandwidth. The doubling in memory bandwidth should translate into a near doubling of tokens/sec/user all else equal, and this is what customers want from Cerebras. Clock speed doubling also drives double the peak theoretical FLOPs and the WSE’s parallel off-wafer I/O, allowing the CS-4 to upgrade to 2.4Tb/s of off-wafer I/O from 1.2Tb/s with CS-3. However, what remains the same is 44GB of SRAM capacity per wafer, as this is determined by the number of SRAM bit cells available on each wafer. so we’ll have to wait for the next generation silicon before we can see any improvement here. This is the main drawback of re-using the same WSE-3 as the low memory capacity per wafer is one of the key tradeoffs inherent with Cerebras’s architecture. CS-4使用与CS-3相同的5nm WSE-3,但Cerebras通过将时钟速度翻倍来获得双倍性能。这得益于向晶圆输送大幅增加的功率,并且由CS-4在供电和冷却技术上的改进所实现。虽然沿用同样的5nm硅听起来并不令人振奋,但Cerebras仍然可以将最重要的指标翻倍:内存带宽。内存带宽的翻倍应该会在其他条件相同的情况下转化为每秒每用户token数接近翻倍,这正是客户希望从Cerebras得到的。时钟速度翻倍也带来了理论峰值FLOPs和WSE并行片外I/O的双倍提升,使CS-4的片外I/O从CS-3的1.2Tb/s升级到2.4Tb/s。然而,每晶圆44GB的SRAM容量保持不变,因为这取决于每片晶圆上可用的SRAM位单元数量。因此,我们必须等到下一代硅片才能看到这方面的改进。这是重复使用WSE-3的主要缺点,因为每晶圆低内存容量是Cerebras架构固有的关键权衡之一。

As we described in our previous article on Cerebras, the wafer has an incredibly unique architecture due to the use of SRAM that makes it well suited for running kernels with low Arithmetic Intensity, such as low-batch size decode. 正如我们在关于Cerebras的前一篇文章中所述,由于使用SRAM,该晶圆具有极其独特的架构,非常适合运行低算术强度(Arithmetic Intensity)的内核,例如低批大小解码。

Source: SemiAnalysis 来源:SemiAnalysis

Network Improvements 网络改进

Beyond the doubling of off-wafer I/O bandwidth, there are further improvements made on off-wafer communications coming from a new Wafer I/O interface, which is an upgraded FPGA card that is used as a NIC to convert Cerebras proprietary I/O to standard ethernet. 除了片外I/O带宽翻倍之外,片外通信还有进一步的改进,这来自于一个新的晶圆I/O接口,该接口是一张升级后的FPGA卡,用作NIC(网络接口卡),将Cerebras专有I/O转换为标准以太网。

From the image below, we can see that there are 2 I/O modules coming off the north and south of the wafer. 从下图可以看出,有2个I/O模块从晶圆的南北两侧引出。

Source: Cerebras 来源:Cerebras

This I/O module is field-upgradeable, so Cerebras can move to new networking standards without redesigning the chassis. This seems like a small change but there are big implications. This makes it easier for the CS-4 to interface with other systems for disaggregated inference setups. Pairing the CS-4 with HBM-based XPUs in a disaggregated attention feed-forward network setup is one way to overcome the CS-4’s low memory capacity, in the same way that Nvidia is positioning the Groq LPUs. 这个I/O模块是现场可升级的,因此Cerebras无需重新设计机箱即可转向新的网络标准。这看起来是个小改动,但意义重大。这使得CS-4更容易与其他系统接口,用于解耦推理设置。在解耦的注意力前馈网络设置中,将CS-4与基于HBM的XPU配对,是克服CS-4低内存容量的一种方法,就像Nvidia定位Groq LPU一样。

We write more about this below. This seems to be designed especially for AWS in mind, who would like to have its EFA NICs on CS-4 to interface with Trainium servers for disaggregated inference. 我们将在下文详细讨论。这似乎是专门为AWS设计的,AWS希望在其CS-4上使用EFA NIC,以便与Trainium服务器接口进行解耦推理。

In addition, latency through the 2 layer fat-tree network (using Arista ethernet switches) is reduced to 3 microseconds from 5 microseconds for CS-3 through a new low latency package processing pipeline. Now, direct wafer to wafer links are also possible rather than going through the switched network which further reduces latency to 2 microseconds. The direct wafer paths are also configurable so this means that the FPGA has switching capability to route data through various wafers. 此外,通过新的低延迟数据包处理流水线,2层胖树网络(使用Arista以太网交换机)的延迟从CS-3的5微秒降低到3微秒。现在,还可以直接进行晶圆到晶圆的链接,而无需经过交换网络,这进一步将延迟降低到2微秒。直接晶圆路径也是可配置的,这意味着FPGA具有交换能力,可以在多个晶圆之间路由数据。

This is a real improvement, but with many Cerebras competitors now quoting all in switch latencies in nanoseconds, “ultrafast” networking is relative and we view it as a modest improvement. We believe this 3µs and bandwidth limitations continues to be a bottleneck that prevents parallelism setups such as EP and ETP where expert layers span multiple wafers. Token dispatch and combine from router to expert is latency sensitive, and the expert imbalance problem coupled with an extra network hop makes pipeline parallelism the only viable solution. 这是一个真正的改进,但随着许多Cerebras竞争对手现在引用全交换机延迟为纳秒级,“超快”网络是相对的,我们认为这是一个适度的改进。我们相信这3微秒的延迟和带宽限制仍然是一个瓶颈,阻碍了诸如EP和ETP之类的并行设置,在这些设置中,专家层跨越多个晶圆。从路由器到专家的token分发和合并对延迟敏感,而专家不均衡问题加上额外的网络跳数,使得流水线并行成为唯一可行的解决方案。

The Backpack Rack 背包式机架

At the system level, the headline change in CS-4 is physical. Cerebras has split the rack into a front half dedicated to power delivery and a rear half dedicated to compute, packaged as modular, pluggable “backpacks.” Each backpack houses a single wafer-scale engine, and a CS-4 rack holds three of them, up from two wafers per rack in CS-3. The cooling infrastructure such as the pump and the heat exchangers are also removed from the rack, as data centers nowadays are built to support fully liquid cooled racks. 在系统层面,CS-4最显著的变化是物理上的。Cerebras将机架分为前半部分(专门用于供电)和后半部分(专门用于计算),并封装为模块化、可插拔的“背包”。每个背包容纳一个晶圆级引擎,一个CS-4机架可容纳三个这样的背包,而CS-3每个机架只有两个晶圆。泵和热交换器等冷却基础设施也从机架中移除,因为如今的数据中心都支持全液冷机架。

Source: Cerebras 来源:Cerebras

The backpack is a vertical enclosure that is split into power, cooling, I/O, and WSE-3 modules independently, which makes the whole system meaningfully simpler to manufacture than CS-3. The WSE-3 wafer sits vertically with the power side facing the front of the rack. The power modules will deliver power via the front side of the wafer, which is the same as CS-3. The cooling of the wafer will be approached from the rear side of the wafer. The I/O modules are attached to the top and bottom edges of the wafer forming a rectangular surface with the wafer. 背包是一个垂直外壳,分别独立划分为电源、冷却、I/O和WSE-3模块,这使得整个系统的制造比CS-3明显简单。WSE-3晶圆垂直放置,电源侧朝向机架正面。电源模块将通过晶圆正面供电,这与CS-3相同。晶圆的冷却将从晶圆背面进行。I/O模块连接到晶圆的顶部和底部边缘,与晶圆形成一个矩形表面。

The backpack design allows a smoother deployment process. Customers can set up the rack with the power modules before simply socketing in the the wafer backpack onto the rack on site. Given the upgrade to three wafer engines per rack running at higher clock speed. One CS-4 rack lands at 125-135kW TDP, which is up around or just short of double the 23kW power draw of a single CS-3. Overall, this means performance/W has at best a slight improvement over CS-3. 背包设计使得部署过程更加顺畅。客户可以先安装好机架的电源模块,然后在现场简单地将晶圆背包插入机架。考虑到每个机架升级为三个晶圆引擎并以更高时钟速度运行,一个CS-4机架的TDP为125-135kW,大约是单个CS-3的23kW功耗的两倍或略低。总体而言,这意味着性能/瓦特相比CS-3最多只有轻微提升。

Source: Cerebras 来源:Cerebras

The big improvement gen-on-gen should be cost. While more cooling and power alone would bring up the BOM, the CS-4’s reduction in components and simpler assembly should be a significant offset to these cost items, and we think the effective BOM per wafer could end up similar to the CS-3. To customers, this means getting nearly double the interactivity and token revenue but at similar TCO which is a very attractive proposition. 代际间最大的改进应该是成本。虽然仅增加冷却和功耗就会提高物料清单(BOM),但CS-4组件减少和组装更简单应该能显著抵消这些成本项目,我们认为每晶圆的有效BOM最终可能与CS-3相近。对客户而言,这意味着以相近的总拥有成本(TCO)获得近乎双倍的交互性和token收入,这是一个非常有吸引力的提议。

Cerebras’s favorite number for CS-4 is 43 PB/s of total on-chip memory bandwidth, which the company markets as roughly 2,000x more memory bandwidth than Nvidia’s Rubin. That’s the number that will lead most coverage of this launch, since on-wafer SRAM bandwidth scales up with the increase in power. Cerebras为CS-4最喜欢宣传的数字是43 PB/s的总片上内存带宽,该公司宣称这大约是Nvidia Rubin内存带宽的2000倍。这个数字将主导此次发布的大部分报道,因为晶圆上SRAM带宽随功耗增加而扩展。

The result is that in spite of 2,000x more memory bandwidth, Cerebras claims a more reasonable interactivity improvement of up to 30x when compared to GPUs, which it’s branding as a new “ultrafast” performance tier. We believe CS-4 will hit near 4,000 tok/sec/user on frontier models, while CS-3 hits 2,000 tok/sec/user. Meanwhile we expect Blackwell GPUs will continue to top out at a theoretical 200 tok/sec/user (which no one actually runs at), and a more realistic 100 tok/sec/user for reasonable amounts of concurrency. That looks like 20-40x more interactivity to us, so why not “up to” 40x faster? Seems fair enough. 结果是,尽管内存带宽高出2000倍,Cerebras声称与GPU相比,交互性提升更合理的数字是高达30倍,并将其定位为新的“超快”性能层级。我们相信CS-4在前沿模型上将达到接近4,000 token/秒/用户,而CS-3为2,000 token/秒/用户。与此同时,我们预计Blackwell GPU将继续在理论上限为200 token/秒/用户(实际上没人会这样运行),而在合理的并发量下更现实的数字是100 token/秒/用户。在我们看来,这相当于20-40倍的交互性提升,那么为什么不说“高达”40倍呢?似乎也很合理。

Parallelism Strategies 并行策略

Because a WSE does not have enough SRAM on wafer to hold an entire model’s weights, Cerebras continues to focus their efforts on pipeline parallel inference with this system. On CS-4, every MoE expert for a given model will sit on a single wafer, interleaved. Pipeline parallelism by default is different than GPUs, where tensor and expert parallelism are most common to get big models to fit in the available HBM. Cerebras has always maintained that using the GPU’s HBM to store weights is slower, more power hungry, and more expensive than doing everything on one wafer. 由于WSE晶圆上的SRAM不足以容纳整个模型的权重,Cerebras继续将精力集中在该系统的流水线并行推理上。在CS-4上,给定模型的每个MoE专家将交错地放在单个晶圆上。默认的流水线并行与GPU不同,GPU最常用张量并行和专家并行来使大型模型适应可用的HBM。Cerebras一直坚持认为,使用GPU的HBM存储权重比在单个晶圆上完成所有操作更慢、更耗电、也更昂贵。

Of course, when comparing a cluster of WSEs to a cluster of GPUs on performance, power consumption, and cost, it really depends what parallelism strategy you choose. GPUs have a very wide range of configuration options (from high throughput/low interactivity configs to high interactivity/low throughput) while the wafer’s range is more modest. Only high interactivity/low throughput is considered. 当然,当比较一组WSE与一组GPU在性能、功耗和成本方面的表现时,这确实取决于你选择哪种并行策略。GPU有非常广泛的配置选项(从高吞吐量/低交互性配置到高交互性/低吞吐量),而晶圆的范围则较为有限。只考虑高交互性/低吞吐量。

In order to compare WSEs directly to GPUs, we are particularly interested in NVIDIA’s release of TileRT, which brings high interactivity/low throughput configs to GPU clusters. We discussed some of this in our TileRT article last week 为了将WSE与GPU直接比较,我们对NVIDIA发布的TileRT特别感兴趣,它将高交互性/低吞吐量配置带到了GPU集群。我们在上周的TileRT文章中讨论过其中一些内容。

Source: InferenceX 来源:InferenceX

Because the wafer itself is 44GB, Cerebras hasn’t yet needed to shard individual experts in a MoE model across multiple wafers, even for frontier models that they run today such as GPT 5.6 Sol. However, due to the demands of long context inference, we expect Cerebras customers such as OpenAI to try and save on the cost of keeping KV Cache on-wafer for long context workloads, and run 5.6 Sol at 256k context window, rather than the full 1M context. 由于晶圆本身为44GB,Cerebras尚未需要将MoE模型中的单个专家分片到多个晶圆上,即使对于他们今天运行的前沿模型(如GPT 5.6 Sol)也是如此。然而,由于长上下文推理的需求,我们预计OpenAI等Cerebras客户将尝试节省长上下文工作负载中在晶圆上保留KV Cache的成本,并以256k上下文窗口运行5.6 Sol,而不是完整的1M上下文。

As we described in our Cerebras IPO article, the cost of supporting long context inference is massive. Most people still understand that the amount of memory required to get one forward pass from a model is proportional to the total amount of parameters in the model. However, it still seems to be a well kept secret that the amount of memory required to hold KV Cache grows proportionally to the number of concurrent users and the average/max size of those users requests (which is dictated by the context window of the model). Running a model with a large context window, and supporting many concurrent users, requires lots of memory capacity. 正如我们在Cerebras IPO文章中所述,支持长上下文推理的成本是巨大的。大多数人仍然理解,从模型获得一次前向传递所需