【文章标题】:AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?

【文章标题】:AgentX - InferenceXv3:CUDA 护城河在智能体推理中是否依然牢固?

【文章正文】:

【文章正文】:

Since the

Claude Code inflection point

in November 2025, long-context, multi-turn agentic workloads have grown rapidly. They now dominate traffic for production inferencing. In April 2026, OpenAI’s Enterprise agentic spending overtook ChatGPT spending.

自 2025 年 11 月的 Claude Code 拐点以来,长上下文、多轮智能体工作负载迅速增长。它们现在主导了生产推理的流量。2026 年 4 月,OpenAI 的企业智能体支出超过了 ChatGPT 支出。

Agentic workflows have decisively taken the baton.

智能体工作流已经果断地接过了接力棒。

Today, we announce AgentX 1.0 - the world’s first fully open source, multi-turn agentic coding inference benchmark at 1 million context, released under Apache 2.0.

今天,我们宣布 AgentX 1.0——世界上第一个完全开源、多轮智能体编码推理基准,支持 100 万上下文,以 Apache 2.0 许可证发布。

Our full dashboard is available here.

我们的完整仪表板可在此处获取。

Source:

SemiAnalysis

来源:

SemiAnalysis

In the past most measured performance based on fixed sequence length prefill and decode workloads, but this is an inaccurate way to measure workloads. Reality is multi-turn, long context, high prefill reuse, with sub agent bursts, KVCache offload, and numerous tool calls. As such we aimed to build the correct way for the industry to measure AI hardware and software performance.

过去,大多数性能测量基于固定序列长度的预填充和解码工作负载,但这是衡量工作负载的不准确方式。现实情况是多轮、长上下文、高预填充重用、子智能体突发、KVCache 卸载以及大量工具调用。因此,我们旨在为行业构建衡量 AI 硬件和软件性能的正确方法。

We have spent more than $3M building this dataset. Today, we open source everything.

我们已花费超过 300 万美元构建此数据集。今天,我们将一切开源。

InferenceXv3 implements AgentX, a new realistic scenario in addition to the existing “fixed sequence length” scenarios (8k1k, 1k1k, 1k8k). It improves the benchmark scenarios by using agentic coding traffic instead of the previous single-turn traffic of 8k input and 1k output tokens.

InferenceXv3 实现了 AgentX,这是一个在现有“固定序列长度”场景(8k1k、1k1k、1k8k)之外新增的真实场景。它通过使用智能体编码流量取代之前 8k 输入和 1k 输出 token 的单轮流量,改进了基准场景。

The full matrix runs on ~2MW of continuously operated compute across over 1000 chips spanning a wide range of SKUs, featuring the MI355X, GB300 NVL72, GB200 NVL72, B300, B200, MI325, MI300X, H200, and RTX Pro Servers. Rubin arrives later this month, and TPUs and Mi455X UALoE72 arrive later this year.

完整矩阵运行在约 2MW 的持续运行计算上,涵盖超过 1000 个芯片,跨越广泛的 SKU,包括 MI355X、GB300 NVL72、GB200 NVL72、B300、B200、MI325、MI300X、H200 和 RTX Pro 服务器。Rubin 将于本月晚些时候到来,TPU 和 Mi455X UALoE72 将于今年晚些时候到来。

Please drop a star if you found our free open source work valuable.

如果您觉得我们的免费开源工作有价值,请给一个 star。

It is great to see amazing performance from both NVIDIA and AMD on agentic workloads. NVIDIA does very good on a lot of frontier models while AMD also does well on some frontier models for specific comparsions.

很高兴看到 NVIDIA 和 AMD 在智能体工作负载上都有出色的表现。NVIDIA 在许多前沿模型上表现非常好,而 AMD 在某些前沿模型的特定比较中也表现良好。

Source:

SemiAnalysis GitHub

来源:

SemiAnalysis GitHub

The most valuable thing AgentX produced in its first months was not the initial results. It was the massive industry impact the benchmark is already having.

AgentX 在最初几个月产生的最有价值的东西不是初步结果,而是该基准已经产生的巨大行业影响。

Over 70+ upstream PRs

for optimizing real world production agentic workloads across vLLM, SGLang, TensorRT-LLM, ATOM, AITER, Dynamo, LMCache, and Mooncake, uses AgentX as the north star benchmark proxy. Most of these optimization improvements are transferable to production traffic. We deep dive into each of these optimizations later in the article.

超过 70 个上游 PR,用于优化跨 vLLM、SGLang、TensorRT-LLM、ATOM、AITER、Dynamo、LMCache 和 Mooncake 的真实生产智能体工作负载,都将 AgentX 作为北极星基准代理。这些优化改进大多可迁移到生产流量中。我们将在本文后面深入探讨每一项优化。

Source: SemiAnalysis

来源:SemiAnalysis

Open source is a core principle for InferenceX and thus, we open more of the stack than most people who use that word. That includes an open frontend, a public database served through an easily consumable REST API

that multiple tier 1 AI lab’s capacity planning teams already consume

, public GitHub Actions CI provenance,

logs

, and accuracy validation on every single point. Crucially, our benchmark configs mainly track

recipes.vllm.ai

and

SGLang cookbook

on upstream images such that we are measuring the performance actual customers are experiencing instead of measuring benchmax’ed images.

开源是 InferenceX 的核心原则,因此我们比大多数使用这个词的人开放更多的技术栈。这包括开放的前端、通过易于使用的 REST API 提供的公共数据库(多个一级 AI 实验室的容量规划团队已经在使用)、公共 GitHub Actions CI 来源、日志以及每个数据点的准确性验证。至关重要的是,我们的基准配置主要跟踪 recipes.vllm.ai 和 SGLang cookbook 的上游镜像,这样我们测量的是实际客户所体验的性能,而不是针对基准测试最大化优化的镜像。

In three to four weeks, we will release an AgentX update article. It will cover further optimizations to agentic workloads, plus updated performance results from AMD and Nvidia. It is important to understand that the profile of agentic workloads is updating fast. InferenceX will continue to move swiftly to benchmark the relevant workloads.

在三到四周内,我们将发布一篇 AgentX 更新文章。它将涵盖对智能体工作负载的进一步优化,以及来自 AMD 和 Nvidia 的最新性能结果。重要的是要理解,智能体工作负载的特征正在快速更新。InferenceX 将继续迅速行动,对相关工作负载进行基准测试。

InferenceX is 100% committed to being open-source - this would not be possible without the contributions and support from our OSS partners. We would like to thank the following people that have made massive contributions to the AgentX 1.0 release:

InferenceX 100% 致力于开源——如果没有我们 OSS 合作伙伴的贡献和支持,这是不可能的。我们要感谢以下为 AgentX 1.0 发布做出巨大贡献的人:

Inferact/vLLM

: Roger Wang, Yifan Qiao, Simon Mo, Jeff Ma, and many others

Inferact/vLLM

:Roger Wang、Yifan Qiao、Simon Mo、Jeff Ma 等等

RedHat/llm-d

: Michael Goin, Robert Shaw, Tyler Michael Smith

RedHat/llm-d

:Michael Goin、Robert Shaw、Tyler Michael Smith

RadixArk/SGLang

: Baizhou Zhang, Yuwei An, Mingyi Lu, and many others

RadixArk/SGLang

:Baizhou Zhang、Yuwei An、Mingyi Lu 等等

LMCache/TensorMesh

: Samuel Shen

LMCache/TensorMesh

:Samuel Shen

Weka

: Callan Fox, ValB

Weka

:Callan Fox、ValB

MoonCake Maintainers:

Teng Ma, Xu Wenjie, Ke Yang

MoonCake 维护者:

Teng Ma、Xu Wenjie、Ke Yang

AMD

: Thomas Wang, HaiShaw, Andy Luo, Seungrok Jung, Chun Fang, Parth Panchal, Bill He, Theresa Shan, Hongxia, Fangzhou, Gilbert Lei, Yanfei Wang, Duyi Wang, Peng Sun, Lingpeng Jin, Simon Danielsson, Xiaohu Guo, Haichen Zhang, Chang Liu, Doug Lehr, Poovaiah Palangappa, and many others in the AMD Shanghai Development Centre

AMD

:Thomas Wang、HaiShaw、Andy Luo、Seungrok Jung、Chun Fang、Parth Panchal、Bill He、Theresa Shan、Hongxia、Fangzhou、Gilbert Lei、Yanfei Wang、Duyi Wang、Peng Sun、Lingpeng Jin、Simon Danielsson、Xiaohu Guo、Haichen Zhang、Chang Liu、Doug Lehr、Poovaiah Palangappa,以及 AMD 上海开发中心的许多其他人

Nvidia

: Xin Li, Anthony Casagrande, Kedar Potdar, Ankur Singh, Ishani Dhanani, Nick Comly, Nvidia Shanghai TensorRT-LLM team, and many others

Nvidia

:Xin Li、Anthony Casagrande、Kedar Potdar、Ankur Singh、Ishani Dhanani、Nick Comly、Nvidia 上海 TensorRT-LLM 团队,以及许多其他人

Anthropic staff, for promptly fixing multiple bugs that made implementing AgentX possible

Anthropic 员工,感谢他们及时修复了多个使 AgentX 实现成为可能的错误

GitHub:

Austen Stone for helping with reliability of GitHub Actions that AgentX uses

GitHub:

Austen Stone,感谢他帮助提高 AgentX 所使用的 GitHub Actions 的可靠性

And many others

以及许多其他人

In addition, we are thankful to all who support our open source InferenceX initiative, including Meta, Microsoft, Oracle, OpenAI, MiniMax, Moonshot Kimi, Alibaba Qwen, and Zhipu GLM.

此外,我们感谢所有支持我们开源 InferenceX 计划的人,包括 Meta、Microsoft、Oracle、OpenAI、MiniMax、Moonshot Kimi、阿里巴巴 Qwen 和智谱 GLM。

Source:

InferenceX

来源:

InferenceX

A Brief Overview of Agentic Workloads

智能体工作负载简要概述

At a high level, an agentic workload is char

在高层次上,智能体工作负载是 char