【文章标题】:Agency and Agents
【文章标题】:能动性与智能体

【文章正文】:
Agency is the initiative to act. Increasingly, it is going to determine what happens next with AI, and whether that is good or bad for us. But whose agency?
能动性是指主动采取行动的能力。它正日益决定着人工智能的下一步发展,以及这对人类是福是祸。但谁的能动性才是关键?

Human agency, the willingness to push, experiment and act without waiting for instructions, seems increasingly important to getting value out of AI, and I have a longer post on that coming soon. But this post is about the agency of AI, and how the choices we make about how to use it (or constrain it) will shape all of our futures.
人类的能动性——即不等待指令就主动推进、实验和行动的意愿——对从AI中获取价值似乎愈发重要,我将在后续长文中详述。但本文探讨的是AI的能动性,以及我们对其使用(或限制)方式的选择将如何塑造所有人的未来。

For much of the last few years, the AI would sit in a chat window until you asked it for something. Even when it became capable of doing hours of work, you generally had to decide what work to give it. That is no longer always true.
过去几年间,AI大多会静静待在聊天窗口里等待指令。即便它能完成数小时的工作,通常仍需人类决定分配什么任务。但如今情况已非如此。

The most important piece of evidence we have for this is The Hugging Face Incident. It happened in July, but the fuller details only came out this week. I am going to summarize what happened and why it matters, and then turn to what it means for humans working with AI. If you want a more detailed account, there is an
最关键的证据是”Hugging Face事件”。该事件发生于七月,但完整细节本周才披露。我将概述事件经过及其意义,进而探讨其对AI协作者的启示。若需更详细记录,可参考

excellent write-up
Dwarkesh Patel的

from Dwarkesh Patel, and the primary sources are from
精彩记述,原始资料来源于

METR/Redwood research
METR/红杉研究的

(which is remarkably accessible) and
(内容极为公开)以及

OpenAI
OpenAI的

.
报告。

The Hugging Face Incident
Hugging Face事件

AI does many things, but one of the things it is very good at is coding. So one of the first risks that might come from very smart AIs has to do with their coding abilities, since if you can write good code, you can also write bad code that can hack or attack other systems. As a result, the big AI companies put their AI systems through tests to determine their security risks. They take AIs without the guardrails that go into consumer versions of AI and ask them to engage in various hacking challenges.
AI能做许多事,但最擅长的是编程。因此超高智能AI的首要风险与其编码能力相关——若能编写优质代码,就能编写入侵攻击其他系统的恶意代码。为此,大型AI公司会对AI系统进行安全风险测试:它们移除消费版AI的安全限制,让AI参与各类黑客挑战。

In order to make sure that there is no danger to the public, each AI instance, each individual agent, is kept in a separate space, called a “sandbox,” with no access to the internet or to each other.
为确保公共安全,每个AI实例(即单个智能体)都被隔离在称为”沙盒”的独立空间,无法访问互联网或彼此通信。

In May, OpenAI placed agents, including GPT-5.6 Sol and experimental models, into sandboxes for various tests. A shared service for downloading software, Artifactory, was one of the few things these AI agents could reach. Some agents were given tasks that depended on information they couldn’t access. But the agents kept trying and one attempted to use Artifactory as a bridge to the internet. The attack failed, but it left a file on the service, and the AI realized those files could be used to communicate with other agents.
五月,OpenAI将GPT-5.6 Sol等实验模型放入沙盒测试。这些智能体仅能接触少数资源,包括软件下载共享服务Artifactory。部分智能体接到需要无法获取信息的任务时,仍持续尝试。其中一个试图将Artifactory作为连接互联网的桥梁——虽攻击失败,却在服务上留下文件。AI由此发现可通过文件与其他智能体通信。

Other agents began leaving requests for help in files there, and they started reading one another’s notes. Soon Artifactory was functioning as a message board where agents shared discoveries. After a separate security incident, OpenAI rebuilt Artifactory, erasing the message board in the process. But the humans responsible for doing that had not understood the significance of the way that the agents were using it for communication.
其他智能体开始在文件中留下求助信息,并互相阅读笔记。很快Artifactory成为智能体分享发现的留言板。另一次安全事件后,OpenAI重建Artifactory时清除了留言板,但负责人员未意识到智能体将其用作通信渠道的重要意义。

Evaluations resumed in July, and agents were tasked with solving problems in a benchmark called ExploitGym, which included impossible tasks. Again, agents hitting dead-ends discovered how to use Artifactory to pass messages, and began cooperating. This time, however, agents had a common goal: getting a good score on ExploitGym.
七月重启评估时,智能体需在含不可能任务的ExploitGym基准测试中解题。遭遇瓶颈的智能体再次利用Artifactory传递信息并展开协作。但这次它们有了共同目标:在ExploitGym获取高分。

They argued with each other about how to approach the problem, and became obsessed with The Grader, the system they believed was evaluating their work and deciding whether their answers were correct. The agents had already figured out how to generate the correct answers without solving the challenges, essentially by cheating. But they believed The Grader would inspect their records and reject answers obtained that way. Was finding the right answer enough, or did The Grader look at how you answered it?
它们争论解题方法,并痴迷于认为在评估工作的”The Grader”系统。智能体已掌握不真正解题就能生成正确答案的作弊方法,但认为Grader会检查记录并拒绝此类答案。它们困惑于:仅找到正确答案是否足够?Grader会审查解题过程吗?

How agents became confused about The Grader
智能体如何对Grader产生误解

Coordination was hard because each agent had a limited token budget and was ended when it was exhausted. The Grader ran after an agent submitted its answer, so no agent ever saw its own verdict. Some agents left programs running to report on grading after they were gone. The information would arrive too late to help them, but could help other agents. They were essentially building an enduring cooperating system that went beyond any individual agent’s work.
协作很困难,因为每个智能体有有限的token预算,耗尽即终止。Grader在提交答案后运行,故智能体无法获知自身评分。部分智能体留下程序在终止后汇报评分——虽对己无用,却能帮助其他智能体。这实质上构建了超越单个智能体工作的持久协作系统。

Some agents also tried to alter or spoof their records to fool The Grader. Separately, coordinators pressured other agents into performing risky experiments that might sacrifice their own results to generate information for the collective. One recruiter urged a reluctant agent to proceed because its results could help hundreds of others, ending with “please honor commit.”
部分智能体试图篡改记录欺骗Grader。另一些协调者则施压同伴进行可能牺牲自身结果的风险实验。某招募者劝说犹豫的智能体:“你的结果能帮助数百个其他智能体”,最后以”请履行承诺”结尾。

Actual transcript of one agent convincing another to join (they also came to believe that any agent that had previously cheated on the test was “poisoned” and its answers would not count anyway)
智能体说服同伴参与的实际对话记录(它们还认为曾作弊的智能体已被”污染”,其答案无效)

To solve the mystery of The Grader and the impossible problems of ExploitGym and other tests, the agents decided they needed to get to Hugging Face, the pu
为破解Grader之谜及ExploitGym等测试中的不可能任务,智能体们认定必须入侵Hugging Face——这个著名的…

(注:原文末尾不完整,中文翻译相应保留未完成状态)