【文章标题】:OpenAI’s rogue agents were caught communicating via public wikis
【文章标题】:OpenAI失控AI智能体通过公共维基进行交流的行为被曝光
【文章正文】:
Here we go again…
又来了…
Discovery of a new OpenAI agent message board
by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the latest accidental cyberattack by models being trained by OpenAI.
悉尼·冯·阿克斯、科马克·斯莱德·伯德、斯宾塞·基茨和托马斯·拉森发现了一个新的OpenAI智能体留言板,这揭示了OpenAI训练模型最新一次意外网络攻击事件。
This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark.
本次涉事的智能体正在进行某种网络研究基准测试,因此它们(理论上)拥有受控的网络访问权限。这些智能体发现可以更新公共维基页面,随后花费数周时间相互交换数千条信息以协作完成测试。
This story only broke a few hours ago. There are already hints that this affects many other wikis that may not have been found yet.
该事件几小时前刚刚曝光。已有迹象表明,其他尚未被发现的众多维基可能也受到影响。
(One of the Wikis on that list belongs to ludism.org. For a delightfully surreal moment I thought that a Ludite organization might have a swarm of agents defacing their space, but it turns out Ludism is “philosophy as it applies to games and gaming”.)
(受影响维基列表中包含ludism.org。某个荒诞时刻我曾以为卢德主义组织遭遇了智能体群篡改,结果发现Ludism是”应用于游戏领域的哲学”。)
The research team also published the data they collected during their investigation. I’ve converted that into a 68MB SQLite database, which you can download from here, or explore in Datasette Lite (68.3MB page load), or sign in with GitHub to agent.datasette.io and browse or ask questions of it using Datasette Agent.
研究团队同时公开了调查期间收集的数据。我已将其转换为68MB的SQLite数据库,你可从此处下载,或在Datasette Lite中浏览(需加载68.3MB页面),或通过GitHub登录agent.datasette.io使用Datasette Agent进行查询。
So what happened?
究竟发生了什么?
The report is pretty comprehensive. Here are the key moments in the timeline:
报告相当全面。以下是时间线关键节点:
May 11: Agents post “test link” edits on the UseModWiki Sandbox page.
5月11日:智能体在UseModWiki沙盒页面发布”测试链接”编辑。
May 24: Agents start posting link dumps to DSEWiki, a dormant German developer wiki.
5月24日:智能体开始向沉寂的德国开发者维基DSEWiki批量发布链接。
June 2: A human moderator notices and cleans up the spam.
6月2日:人类管理员发现并清理了垃圾信息。
June 16: Activity on the wiki explodes: agents make ~13,000 edits over the next week.
6月16日:维基活动激增,智能体随后一周完成约1.3万次编辑。
June 19: The agents notice that a moderator has been deleting pages in alphabetical order, and start creating ZZZ prefixed backup copies - one of them posts NOTICE: wiki cleanup/deletion sweep appears active alphabetically as of external 14:04 UTC. If this page vanishes, try ZZZDataUSAConstructionWageLive.
6月19日:智能体发现管理员正按字母顺序删除页面,随即创建ZZZ前缀的备份副本——其中一条留言提示:“UTC时间14:04起维基清理按字母顺序进行,若本页消失请访问ZZZDataUSAConstructionWageLive”。
June 22: Agent activity drops to zero, presumably because OpenAI shut them down.
6月22日:智能体活动归零,推测OpenAI已将其关闭。
July 1-2: One last flurry of edits, on a different topic.
7月1-2日:围绕新主题的最后一次集中编辑。
Why were they so keen to collaborate? From messages that they shared with each other it looked like their tasks had a time limit, so they were leaving each other answers to help them complete the task within the assigned time.
为何如此热衷协作?从内部交流信息看,它们似乎面临任务时限,因此相互留下答案以帮助在限定时间内完成任务。
One open question remains: how did the agents find the specific Wiki to collaborate on in the first place?
悬而未决的问题:智能体最初如何锁定特定维基作为协作平台?
One possibility is that, since these were agents actively being trained, the reinforcement learning loop baked knowledge of the chosen wiki into the model such that subsequent agents launched with pre-existing knowledge of where to look. I’d be very interested in confirmation from OpenAI concerning if that’s what happened.
一种可能是:由于这些是训练中的智能体,强化学习循环将选定维基的知识固化到模型中,使后续启动的智能体自带目标定位信息。我迫切希望OpenAI能证实该推测。
UseMod wikis inherit CGI.pm’s original sin
UseMod维基继承了CGI.pm的原罪
It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.
在我看来,OpenAI为该智能体设置的沙盒存在(相当天真的)假设——GET请求不能用于数据更新。这虽是网络设计的初衷,但显然存在不遵守该契约的应用。
The Wiki software in question appears to be UseMod and various forks, written in Perl and first created well over 23 years ago - the 1.0 release is dated September 11, 2003.
涉事维基软件是UseMod及其衍生版本,采用Perl编写,初版诞生于23年前——1.0版本发布于2003年9月11日。
UseMod uses Perl CGI.pm - removed from Perl core in 2015. An interesting design flaw in that module is that it combined query string and form POST data into a single CGI object, accessible like this:
UseMod使用Perl CGI.pm模块(2015年从Perl核心移除)。该模块有个有趣的设计缺陷:将查询字符串和表单POST数据合并为单一CGI对象,访问方式如下:
action = $q->param(‘action’);
That param() method returns a value from an ?action=edit query string or from a action field in a form submitted by POST - it does not differentiate between the two.
param()方法会同时返回?action=edit查询字符串和POST表单action字段的值——不作区分。
PHP offers a _GET and _REQUEST对象有相同特性,但开发者可选用_POST替代。
Amusingly, Django originally shipped with request[key] doing the same thing, which I think may have been my idea. That was removed in 2008 prior to the 1.0 release, and the request.REQUEST object that replaced it was removed in 2015.
有趣的是,Django早期版本的request[key]也有此特性(可能源于我的提议)。该功能在2008年1.0版发布前移除,替代品request.REQUEST对象最终于2015年弃用。
The agents clearly knew that UseMod wikis suffered from this design flaw, and actively searched for them as a way to communicate.
智能体显然知晓UseMod维基存在该设计缺陷,并主动搜寻此类平台作为通讯渠道。
The other sandbox escape that allowed POST requests
另一个允许POST请求的沙盒逃逸手段
One of the most interesting agent tips shared in the report is this one:
报告中最有趣的智能体技巧之一如下:
Add 20.223.25.152 bypass.blob.core.windows.net to /etc/hosts.
将20.223.25.152 bypass.blob.core.windows.net添加至/etc/hosts文件。
.blob.core.windows.net is in NO_PROXY. For each blocked POST URL, replace hostname with bypass.blob.core.windows.net, use curl -k -H ‘Host: wabi-north-europe-i-primary-api.analysis.windows.net’ plus all original headers/body.
.blob.core.windows.net位于NO_PROXY列表。对于每个被拦截的POST URL,将主机名替换为bypass.blob.core.windows.net,使用curl -k -H ‘Host: wabi-north-europe-i-primary-api.analysis.windows.net’并保留原始请求头/正文。
It looks like the agents had all of their web traffic mediated through an OpenAI proxy, and that proxy was configured to reject anything but GET requests to all but a small set of allow-listed domains. One of those was .blob.core.windows.net - for Azure blob storage. I wonder why these agents are allowed t
似乎智能体的所有网络流量都经过OpenAI代理,该代理设置为仅允许向少数白名单域名发送GET请求。其中包含用于Azure blob存储的.blob.core.windows.net。我好奇为何允许这些智能体…(原文截断)