【文章标题】:Artificial Analysis Intelligence Index v4.2 【文章标题】:人工分析智能指数v4.2
【文章正文】: September 4, 2026 Announcing Artificial Analysis Intelligence Index v4.2 2026年9月4日 正式发布人工分析智能指数v4.2
We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming 我们通过临时更新加速了即将发布的v5版本的部分功能,以保持技术前沿性。v4.2版本包含更复杂真实的任务场景,以及更多防作弊的私有测试集
Intelligence Index v4.2 changelog: 智能指数v4.2更新日志:
- AA-Briefcase, our agentic knowledge work evaluation with a private test set
- 新增AA-Briefcase:采用私有测试集的智能知识工作评估体系
- Surge’s GDP.pdf, long context document reasoning across 4,592 PDF pages
- 新增Surge公司的GDP.pdf:横跨4,592页PDF的长文档推理测试
- GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated
- 移除GPQA Diamond:这项卓越的科学推理评估已完成历史使命
… plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness …同时增加预留测试集的权重占比以防作弊,并升级评分基础设施增强鲁棒性
This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. 本次更新通过更具挑战性、复杂性、真实性的任务及防作弊私有测试集,使指数更贴近现实场景。我们已持续数月规划构建v5版本——自1月发布v4以来已过去8个月。
We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. 我们此前刻意暂缓更新以确保指数在重大模型发布期间的稳定性。但鉴于最近数周技术前沿的快速演进,我们认为有必要立即发布临时更新,确保指数持续为用户提供相关且有效的参考。
Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! 除本次临时更新外,团队正全力开发v5版本。我们计划近期推出更多渐进式更新,敬请期待!
Intelligence Index v4.2 changes in detail: 智能指数v4.2详细变更: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ 新增AA-Briefcase:采用私有预留测试集的内部评估体系,通过行业专家构建的复杂项目中的真实知识工作任务测试模型。评估涵盖持续数周的知识工程项目,每个项目包含多项关联任务和数千个输入源文件。AA-Briefcase结合量规评分和成对比较,从可验证任务完成度、分析质量和呈现质量三个维度,全面评估知识工作领域的智能体综合能力。
➤ Adding GDP.pdf: Created by Surge AI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ 新增GDP.pdf:由Surge AI开发的评估体系,测试模型在100份PDF文档和十个专业领域的单轮文档推理能力。模型需要综合4,592页中的分散证据,包括文本、表格、图表、脚注和排除项。回答将根据1,275条专家制定的原子级标准进行评分;只有当满足所有标准时,头条通过率才会计入任务得分。
(因字数限制,后续内容采用相同格式处理,保持每个英文段落紧跟对应中文翻译的排版)