【文章标题】:Terminal-Bench-Science:评估AI智能体在科研工作流中的表现
【Terminal-Bench-Science:在研究者实际工作流程中评估AI智能体。由科学家(而非模型开发者或数据供应商)为AI的科学能力设定标准】

Terminal-Bench-Science 0.1
Terminal-Bench-Science evaluates AI agents on workflows from researchers’ own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI.
Terminal-Bench-Science基于研究者实际工作流程评估AI智能体。由科学家(而非模型开发者或数据供应商)为AI的科学能力设定标准。

Terminal-Bench-Science is a benchmark led by researchers at Stanford University and built by the team behind Terminal-Bench in collaboration with domain experts from a range of scientific disciplines and research institutions around the world. It measures the AI agent capabilities through a diverse set of challenging, expert-curated workflows drawn from scientific research.
该基准由斯坦福大学研究人员领导,Terminal-Bench团队联合全球多学科领域专家共同构建,通过精选科研中具有挑战性的工作流来评估AI智能体能力。

Terminal-Bench-Science is a continuous benchmark that evolves alongside frontier AI, creating a feedback loop between scientific needs and AI development. Our first release includes 70 tasks from the life, physical, Earth, mathematical, and engineering sciences. The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1.
这是一个与前沿AI共同进化的持续基准,在科学需求与AI发展间形成闭环。首版包含生命、物理、地球、数学和工程科学领域的70项任务,表现最佳的Claude Opus 5模型解决率为30%。

Overview
概述

While Terminal-Bench has driven progress in AI agents for software engineering, Terminal-Bench-Science brings the same ambition to science. Our goal is to drive the development of agents with scientific capabilities that make them useful research assistants.
正如Terminal-Bench推动软件工程AI智能体发展,Terminal-Bench-Science对科学领域怀抱同等雄心。我们的目标是培育具备科研能力的智能体,使其成为实用研究助手。

These agents should execute technically demanding and time-consuming workflows, freeing scientists to focus more of their time on the parts of science where human judgment matters most: defining research questions, forming hypotheses, interpreting and validating results, and communicating findings.
这些智能体应能执行技术性强且耗时的工作流,让科学家集中精力于需要人类判断的核心环节:界定研究问题、形成假设、解读验证结果及交流发现。

Achieving this requires benchmarks that reflect real scientific practice, provide verifiable evidence of capability, and evolve alongside the AI frontier.
实现这一愿景需要满足三大条件的基准:反映真实科研实践、提供可验证能力证据、与AI前沿同步进化。

We need benchmarks drawn from real scientific workflows.
我们需要源自真实科研工作流的基准

Scientific capability should be evaluated on real research practice, not textbook questions or standardized exercises, contributed by practicing scientists themselves. Terminal-Bench-Science gives scientists across domains a direct voice and a shared platform to set the bar for AI progress on the problems they care about.
科学能力评估应基于实际研究实践(而非教科书问题或标准化练习),由一线科学家亲自贡献。该基准为跨领域科学家提供直抒己见的平台,为其关注的问题设定AI进步标准。

We need verifiable evidence of scientific capability.
我们需要可验证的科学能力证据

Without reliable evaluation, we cannot tell whether agent capabilities are improving or where their limitations remain. Terminal-Bench-Science evaluates agents in realistic environments and grades concrete artifacts such as analyses, simulations, proofs, code, and data products with reproducible, task-specific tests.
缺乏可靠评估就无法判断智能体能力进展。该基准在真实环境中通过可复现的专项测试,对分析、模拟、证明、代码和数据产品等具体产出进行评分。

We need a benchmark that keeps pace with the frontier.
我们需要与前沿同步的基准

Terminal-Bench-Science is a continuous benchmark that evolves alongside the AI frontier. Through regular releases, scientists can contribute new workflows, improve existing tasks, and create a feedback loop between scientific needs and AI development.
这是一个与AI前沿共同进化的持续基准。通过定期更新,科学家可贡献新工作流、改进现有任务,形成科研需求与AI发展间的闭环。

Tasks
任务体系

Terminal-Bench-Science 0.1 includes 70 tasks across the life, physical, Earth, mathematical, and engineering sciences. Tasks span scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning.
首版涵盖生命、物理、地球、数学和工程科学的70项任务,涉及科研数据分析、统计推断、模拟仿真、优化计算、定理证明、图像重建、信号处理、反问题求解、传感器校准、模型拟合、分类及科学机器学习等领域。

Tasks are contributed by researchers through an open process on GitHub, with discussion and feedback in the tb-science channel on Discord. Contributions begin as proposals, where reviewers discuss each idea, leave feedback, and approve those that look like a strong fit: scientifically grounded workflows worth measuring in the benchmark.
研究者通过GitHub开放流程提交任务,在Discord的#tb-science频道讨论。提案经评审讨论后,筛选出具有科学基础、值得纳入基准的工作流。

Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1. That selectivity reflects how difficult it is to create tasks that are scientifically interesting, challenging for frontier agents, and sufficiently well specified for rigorous evaluation.
920份提案中仅70项入选首版,这种严苛筛选反映出创建兼具科学趣味性、前沿挑战性和可严格评估标准的任务之难。

Results
评估结果

Terminal-Bench-Science 0.1 leaves substantial room for progress on AI agents for scientific research. Each evaluated model ran thr…
Terminal-Bench-Science 0.1显示科研AI智能体仍有巨大提升空间。每个受评模型运行三…

(注:原文末尾截断,中文翻译相应保留未完成状态)