【文章标题】:GLM-5.3: How Chinese labs keep stride with the frontier
【文章标题】:GLM-5.3:中国实验室如何与前沿保持同步
【文章正文】: Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out.
【文章正文】: 杂务说明:我正在旅途中,因此无法为本篇文章录制语音。编者注——在邮件发出后,我补充了关于中国数据行业的第5点内容。
Today, Z.ai
announced
their GLM-5.3 model, currently only available in the coding plan, coming soon to their API and in two weeks’ time to Hugging Face (open weights). This model looks exceptional, with a somewhat astounding increase in scores. On many benchmarks the model has surpassed Moonshot AI’s Kimi K3 and on some it’s surpassed Claude Fable 5 or GPT-5.6-Sol.
今天,Z.ai
发布了
他们的GLM-5.3模型,目前仅在编程方案中可用,即将上线其API,并将在两周后登陆Hugging Face(开放权重)。这款模型表现非凡,分数提升令人震惊。在许多基准测试中,该模型已超越月之暗面的Kimi K3,在部分测试中甚至超过了Claude Fable 5或GPT-5.6-Sol。
Here’s a more complete comparison:
以下是更完整的对比:
This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3! The Z.ai blog post is rather straightforward, and starts with a bold sentence:
这使该模型基本上处于智能体编码基准的前沿,而参数量仅约7500亿——只有Kimi K3的三分之一!Z.ai的博客文章相当直白,开篇就是一句大胆的话:
Scaling post-training is all we did for GLM-5.3.
扩展后训练就是我们为GLM-5.3所做的一切。
GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece. Following this release there have been a lot of discussions wondering how China can keep up so well? How can such a small model be matching the leading public American models? Are these results real?
GLM-5.3与GLM-5.2使用相同的基础模型,但后训练大幅扩展。冒着一个过度简化的风险来说,与Kimi相比,Z.ai似乎在後训练方面具有优势,而Kimi更像是预训练的杰作。此次发布之后,引发了大量讨论:中国如何能保持如此出色的跟进速度?这么小的模型怎能与领先的美国公开模型匹敌?这些结果是真的吗?
Subscribe now
立即订阅
The simplest explanation is that Z.ai is very good at what they do – it’s worth recalling that they’ve been working on this line of models longer than almost anyone in the industry. Here’s a brief history of the GLM models.
最简单的解释是,Z.ai非常擅长他们所做的事情——值得回顾的是,他们研究这一系列模型的时间比业内几乎任何一家公司都要长。以下是GLM模型的简要历史。
Zhipu AI Founded
– 2019
智谱AI成立
——2019年
GLM
(General Language Model) —
March 2021
— released by
THUDM
, Tsinghua University’s Data Mining / Knowledge Engineering group.
Weights
GLM
(通用语言模型)——
2021年3月
——由
THUDM
(清华大学数据挖掘/知识工程研究组)发布。
权重
GLM-130B
—
August 2022
— Scaled version.
Technical report for GLM-130B through GLM-4
—
Weights
GLM-130B
——2022年8月——扩展版本。
GLM-130B至GLM-4的技术报告
——
权重
ChatGLM
—
March 14, 2023
— first chat version.
Weights
ChatGLM
——2023年3月14日——首个聊天版本。
权重
ChatGLM2
—
June 25, 2023
—
Weights
ChatGLM2
——2023年6月25日——
权重
ChatGLM3
—
October 27, 2023
—
Weights
ChatGLM3
——2023年10月27日——
权重
GLM-4
—
January 16, 2024
— rebranded as just GLM; open-weight GLM-4-9B followed in June.
Weights
GLM-4
——2024年1月16日——更名为GLM;随后于6月推出开放权重的GLM-4-9B。
权重
GLM-5
—
February 11, 2026
— latest major generation.
Weights
GLM-5
——2026年2月11日——最新主要世代。
权重
GLM 5.2, released on June 22 of this year, was a big deal
– weeks after the release, I regularly heard from AI researchers I know who still used the model due to its speed (some deploy the model on internal clusters for faster speeds than public offerings) and simplicity (as a model with no rollbacks, etc., when working on frontier AI systems). GLM-5.2 altogether stood up to the hype.
今年6月22日发布的GLM 5.2是一个大事件
——发布数周后,我经常听到我认识的AI研究人员提到,他们仍在用这个模型,因为它速度快(有些人将模型部署在内部集群上,以获得比公开服务更快的速度)且简洁(作为一个没有回滚等机制的模型,在前沿AI系统上工作时尤为方便)。GLM-5.2完全经受住了炒作。
I’ve been going through some of the same denial myself, thinking “how do they keep doing this?
Surely
the models aren’t as good as they look.” There’s something a bit off-putting with how the American companies have such a commanding resource lead, but can’t seem to pull away in capabilities. The common answer is distillation, which I’ve
written
at length
about, but I deem not to be the major factor. On that note, there was a
recent paper
that showed simple methods for extracting the reasoning traces from frontier models – this is the sort of thing that Chinese labs could definitely use at scale. I’m confused why the labs in the U.S. haven’t patched this behavior faster; instead they’re running to the government asking for policy help. It doesn’t add up for me.
我自己也曾经历同样的否认心态,想着“他们怎么总能做到?
肯定
模型没有看起来那么好。”令人有些不适的是,美国公司在资源上拥有压倒性领先优势,却在能力上似乎无法拉开差距。常见的解释是蒸馏,我曾在
文章
中
详细探讨过,但我认为这不是主要因素。关于这一点,最近有一篇
论文
展示了从前沿模型中提取推理轨迹的简单方法——这正是中国实验室绝对可以大规模使用的东西。我不明白为什么美国实验室没有更快地修补这种行为;相反,他们跑去向政府寻求政策帮助。这对我来说说不通。
Z.ai’s blog is direct and matches with an RL-dominated training regime. They say they used “
more environments, more diverse tasks, and more compute spent training on them.” One does not simply “distill” RL environments, infrastructure to run them at scale, or algorithms to mix them together effectively.
Z.ai的博客直截了当,与其以强化学习为主导的训练体系相吻合。他们表示使用了“
更多的环境、更多样化的任务,以及投入更多算力进行训练。”人们无法简单地“蒸馏”强化学习环境、大规模运行这些环境的基础设施,或有效混合它们的算法。
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
Interconnects AI是一份由读者支持的刊物。请考虑成为订阅者。
So, how do the Chinese labs do it if not distillation? Are they benchmaxxing? An accepted definition of benchmaxxing is focusing the model on the test sets, such that the real-world performance meaningfully differs from the on-paper scores. The determining factors are much more big picture than technical (yes, the technical details definitely matter, but are harder to differentiate from lab to lab):
那么,如果不是蒸馏,中国实验室是如何做到的?他们是在刷基准测试吗?对“刷基准测试”的一个公认定义是让模型专注于测试集,使得实际表现与纸面分数存在显著差异。决定性因素更多是宏观层面的而非技术层面的(是的,技术细节确实重要,但很难区分不同实验室之间的差异):
The time to release for Z.ai is likely days, not months as with OpenAI or Anthropic.
Z.ai的发布周期可能以天计算,而非OpenAI或Anthropic那样的以月计算。
It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI. Still, these American companies
tend to take months to release their models to the public
, which massively flatters the Chinese labs in adoption decisions at the frontier. To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks (SpaceXAI is likely far closer to the Chinese labs here). With the pace of progress being so fast, this is likely the largest determining factor of why Chinese labs stay at the frontier. This, so far, has been economically acceptable for the American labs, as they’ve still had massive demand for their models.
OpenAI和Anthropic很可能拥有远比Z.ai和月之暗面更好的内部模型。然而,这些美国公司
往往需要数月时间才能将模型公开发布
,这在前沿采用决策上极大地有利于中国实验室。简而言之——中国实验室利用美国实验室进行发布前测试的所有时间,持续在基准测试上爬山优化(SpaceXAI在这方面可能与中国实验室更为接近)。鉴于进步速度如此之快,这可能是中国实验室能保持前沿地位的最大决定因素。到目前为止,这对美国实验室在经济上尚可接受,因为他们的模型仍有巨大需求。
As model self-improvement loops ramp up within the labs building LLMs, if any of these feedback loops require user data, this faster release cycle could massively favor the Chinese labs, giving their offerings longer lifespans before the next vastly superior model comes out, undercutting demand for their models.
随着构建大语言模型的实验室中模型自我改进循环不断加速,如果这些反馈循环中有任何一环需要用户数据,这种更快的发布周期可能会极大地利好中国实验室,使他们的产品在下一个更强大的模型问世之前拥有更长的生命周期,从而削弱对其模型的需求。
These are very clearly the race dynamics that many in the industry worry about. With so many labs building frontier models in the envelope of leading capabilities, it is hard to see this abating in the near future.
这些显然就是业内许多人担心的竞赛动态。如此多的实验室在领先能力的范围内构建前沿模型,很难看到这种情况在近期内有所缓解。
Yes, Z.ai probably cares slightly more about public benchmarks than OpenAI or Anthropic.
是的,Z.ai可能比OpenAI或Anthropic更在意公开基准测试。
These benchmarks, e.g. scoring highly on the Artificial Analysis Intelligence Index, or similar aggregators, have a very direct impact on their stock price. They in many ways need to do this to keep raising capital and maintain team morale, as being the scrappy underdog matching American giants is a wonderful story.
这些基准测试,例如在Artificial Analysis Intelligence Index或类似聚合平台上获得高分,对他们的股价有非常直接的影响。他们在许多方面需要这样做来持续融资和维持团队士气,因为作为不畏强手的挑战者与美国的巨头们比肩,是一个精彩的故事。
Subtle benchmaxxing does not need to come out of desperation or any similar pressures. It’s the industry standard across a remarkable number of labs. Many companies’ data acquisition strategy is to buy data on the benchmarks they’re behind on.
隐性的刷基准测试并不一定源于绝望或类似压力。在数量众多的实验室中,这已是行业标准。许多公司的数据采集策略就是购买他们在基准测试上落后领域的相关数据。
Z.ai is not benchmaxxing to the point where GLM-5.3 is fried
(at least not intentionally, and they’ll check for it). Every lab is dealing with the rough edges of scaling RL right now. Anthropic’s Opus 5 and Sonnet 5 models have very mixed reputations, despite the incredible benchmark scores. Everyone in the industry is in the same boat, so some model weights end up being easier to use than others, but the benchmark scores in their release blogs are the real deal.
Z.ai并没有刷基准刷到让GLM-5.3“烤焦”的程度
(至少不是故意的,而且他们会去核查)。目前每个实验室都在应对扩展强化学习带来的各种粗糙问题。Anthropic的Opus 5和Sonnet 5模型尽管在基准测试上取得了惊人分数,但口碑却褒贬不一。业内所有人都面临同样的处境,所以有些模型的权重最终比其他模型更容易使用,但他们在发布博客中公布的基准分数是实打实的。
GLM-5.3 is likely a narrower model than Claude Fable or GPT Sol.
GLM-5.3可能比Claude Fable或GPT Sol更加窄专。
When GPT-5.2 was released, it had mixed reviews outside of agentic coding. At the same time, OpenAI and Anthropic support very large businesses with countless use-cases for their models. This is a benefit of being a company earlier in their adoption curve – you can target the most valuable use-cases. Within post-training, caring about a bit less will make assembling the final model
far
easier.
GPT-5.2发布时,除智能体编码之外的评价褒贬不一。与此同时,OpenAI和Anthropic支撑着非常庞大的业务,其模型拥有无数用例。这是处于采用曲线更早期阶段的公司的一个优势——你可以瞄准最有价值的用例。在后训练中,少关注一些方面会让最终模型的组装
容易得多
。
I’m overstating this a bit, as Z.ai
reportedly reached $1B of ARR
on the back of a strong on-premises deployment business.
我有点夸大了这一点,因为Z.ai
据报道已实现10亿美元的年经常性收入(ARR)
,这得益于其强大的本地部署业务。
Similarly, the flagship GLM models have not had visual capabilities. Being text-only definitely helps Z.ai get more competitive scores, but it is a more competitive space. On the other side of things are models like
Inkling-Small
, which is designed to be omnimodal.
同样,旗舰GLM模型一直没有视觉能力。纯文本模式确实帮助Z.ai获得了更具竞争力的分数,但这也是一个竞争更激烈的领域。而另一端的模型则像
Inkling-Small
,它被设计为全模态模型。
(ADDED)
The RL data industry is taking off in China
. Many
sources
and rumor-mills we’re following have been mentioning how the data industry is taking off in China — very much driven by American data companies selling to Chinese model labs. This could look like Chinese labs buying many of the same RL environments that are used by American frontier labs, and releasing the downstream RL’d model sooner. We still have large error bars on the scale and impact of this market, but it is certainly becoming important.
(补充)
中国的强化学习数据产业正在腾飞
。我们关注的许多
信源
和传闻渠道都在提到中国的数据产业如何蓬勃发展——很大程度上是由美国数据公司向中国模型实验室销售所推动的。这可能表现为中国实验室购买美国前沿实验室所使用的许多相同的强化学习环境,并更早地发布下游经过强化学习的模型。我们对该市场的规模和影响仍存在很大的误差范围,但它无疑正变得重要。
Z.ai is an extremely skilled LLM organization – one that is likely far more compute efficient than OpenAI / Anthropic.
Z.ai是一家技艺极其精湛的大语言模型组织——其算力利用效率可能远高于OpenAI/Anthropic。
This needs repeating. These folks are very good at what they do. The company has very close ties to Tsinghua University, which is home to many of the best Chinese computer scientists. This abundant, eager talent pool is as central to their success as it is for any Western counterpart.
这一点值得重复强调。这些人非常擅长他们所做的事情。该公司与清华大学有着非常紧密的联系,清华大学汇聚了许多中国最优秀的计算机科学家。这个丰富而充满热情的人才库对他们的成功至关重要,正如对任何西方同行一样。
Altogether, it seems like a perfectly good strategy they’re executing with the GLM line of models. Congrats on the release! I’m excited for the weights to be out so I can do more extended testing (I tend to use American open-weight inference services like Fireworks or Baseten).
总的来说,他们在GLM模型系列上执行的策略似乎非常出色。恭喜发布!我很期待权重开放,这样我就能进行更深入的测试(我倾向于使用Fireworks或Baseten等美国开放权重推理服务)。
Leave a comment
留下评论
This is another step towards the inevitable proliferation of very strong cyber capabilities across the economy. Z.ai has acknowledged this,
saying
:
这是向着极其强大的网络能力在整个经济中不可避免扩散的又一步。Z.ai已经承认了这一点,
表示
:
GLM-5.3 is our most capable model to date for cybersecurity tasks. It delivers substantial improvements in vulnerability discovery, exploit analysis, and complex multistep security tasks. These capabilities can help defenders identify weaknesses earlier, validate risks, and accelerate remediation.
GLM-5.3是我们迄今为止在网络安全任务上最强大的模型。它在漏洞发现、漏洞利用分析和复杂多步安全任务方面实现了显著提升。这些能力可以帮助防御者更早识别弱点、验证风险并加速修复。
They also create clear dual-use risks. We are therefore taking a staged approach to release. Selected security partners will first evaluate GLM-5.3 in controlled settings. Broader access and API availability will follow. Once the necessary safety evaluations and release preparations are complete, we will publish GLM-5.3’s complete model weights.
这些能力同时也带来了明确的双重用途风险。因此,我们采取分阶段发布的方式。选定的安全合作伙伴将首先在受控环境中评估GLM-5.3。随后将提供更广泛的访问和API可用性。一旦必要的安全评估和发布准备工作完成,我们将发布GLM-5.3的完整模型权重。
They go on to acknowledge how they’re monitoring inference on their platforms via a request classifier and chain of thought monitoring (on top of model alignment). The devil is in the details here, and it is unclear the level of execution every AI lab will have here. The capability diffusion is determined by the lowest common denominator.
他们接着承认如何通过请求分类器和思维链监控(在模型对齐之上)来监控其平台上的推理。细节决定成败,目前尚不清楚每个AI实验室在这一层面的执行水平如何。能力的扩散取决于最薄弱的环节。
At the end of the day