【文章标题】:Autonomy and Innovation 【文章标题】:自主性与创新
【文章正文】: Listen to this
post
:
Log in to listen
收听本文
文章
:
登录后收听
While not every Western followed the cliché, by the 1930s cowboy serials had landed on a consistent visual cue: the hero of the show wore a white hat, and the villain wore a black one. At the end of the day, however, they both were cowboys with cowboy hats.
虽然并非每一部西部片都遵循这一陈规,但到20世纪30年代,牛仔系列片已经形成了一致的视觉标志:剧中的英雄戴白帽,反派戴黑帽。然而归根结底,他们都是戴着牛仔帽的牛仔。
Westerns aren’t much of a cultural touchpoint anymore, but the “white hat” and “black hat” nomenclature is very relevant in tech: hackers who are focused on patching vulnerabilities and protecting software are “white hat hackers”, while hackers who are focused on exploiting vulnerabilities for malicious reasons are “black hat hackers”. Of course this can very quickly become complicated: governments might employ hackers to break into enemy software installations — are they white hats or black hats? Or consider bug bounty programs, wherein large software companies pay bug bounties to hackers who find and report vulnerabilities; it’s basically using money to incentivize would-be black hat hackers to be white hat hackers.
西部片如今已不再是重要的文化触点,但“白帽”和“黑帽”这一术语在科技领域却非常相关:专注于修补漏洞和保护软件的黑客是“白帽黑客”,而专注于出于恶意目的利用漏洞的黑客则是“黑帽黑客”。当然,这很快就会变得复杂:政府可能会雇用黑客入侵敌方的软件设施——他们是白帽还是黑帽?或者想想漏洞赏金计划,大型软件公司向发现并报告漏洞的黑客支付赏金;这基本上是用金钱激励潜在的黑帽黑客成为白帽黑客。
The actual takeaway is that all of this complexity is overwrought: just as a cowboy is a cowboy, a hacker is a hacker; the hat is not a statement of capability, but rather intentions, and those intentions are shaped by incentives. The best way to attack infrastructure is to find a vulnerability and exploit it; the best way to defend infrastructure is to find a vulnerability and patch it. It’s all the same skillset.
真正的要点在于,所有这些复杂性都是过度解读:正如牛仔就是牛仔,黑客就是黑客;帽子并不代表能力,而是代表意图,而这些意图是由激励塑造的。攻击基础设施的最佳方式是找到漏洞并利用它;防御基础设施的最佳方式是找到漏洞并修补它。这都是同一套技能。
This delineation between capability and intent and incentive is critical when it comes to AI. At the end of last month’s Article
Who’s Afraid of Chinese Models
, I discussed a mysterious attack that model host Hugging Face had just endured, which they were only able to fight off with the help of open weight Chinese models, and wrote:
在人工智能领域,能力、意图和激励之间的这种区分至关重要。在上月末的文章《谁害怕中国模型?》中,我讨论了模型托管平台 Hugging Face 刚刚遭受的一次神秘攻击,他们只能借助开放权重的中国模型才得以抵御,并写道:
It’s difficult to overstate how wrong-headed the Trump administration’s panicked response to Anthropic’s release of Fable was, particularly since it exacerbated Anthropic’s worst tendencies in terms of assuming only they can be trusted with powerful AI. In a world with only one AI, it might make sense to reserve the most powerful cybersecurity capabilities for the U.S. government and trusted allies; however, that’s not the world we live in.
特朗普政府对 Anthropic 发布 Fable 的恐慌反应有多么错误,怎么说都不为过,尤其是这加剧了 Anthropic 最糟糕的倾向,即认为只有他们才能被信任拥有强大的 AI。在一个只有一种 AI 的世界里,将最强大的网络安全能力保留给美国政府和可信盟友或许说得通;然而,我们生活的世界并非如此。
There are and will be models eminently capable of mounting cybersecurity attacks on existing infrastructure, and those models will be — already are — widely available. The best defense — the only viable defense, in fact — will be to make sure defenders have access to the best models as well. Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!
现在和将来都会有模型完全有能力对现有基础设施发动网络安全攻击,而且这些模型将——实际上已经——广泛可用。最好的防御——事实上唯一可行的防御——将是确保防御者也能使用最好的模型。目前,由于特朗普政府的指令,防御者实际上被禁止在网络安全领域使用 Fable 或 Sol;这意味着最好的替代方案是使用来自一个多年来一直试图削弱我们网络防御的国家的模型。这太荒谬了!
The point is the one I made in the introduction: when it comes to cybersecurity, the capability that is necessary for good defense is the exact same capability that is necessary for good offense; the color of the hat is a matter of who is actually prompting the AI.
这一点正是我在引言中提出的:在网络安全方面,良好防御所需的能力与良好进攻所需的能力完全相同;帽子的颜色取决于实际上是谁在向 AI 发出提示。
And, sometimes, not even that is clear: it turns out that the entity that hacked Hugging Face was actually OpenAI, as a series of unconstrained agents being evaluated for their cybersecurity capabilities found and exploited a bug in the package manager in their sandbox; that package manager had Internet access and a sufficiently writeable file system such that the agents could communicate with each other over time. The entire chain of vulnerability discovery and exploit creation culminated in the so-called “Hugging Face incident”.
而且,有时连这一点也不清楚:事实证明,入侵 Hugging Face 的实体实际上是 OpenAI,因为一系列正在接受网络安全能力评估的、不受约束的智能体在其沙箱中发现了包管理器的一个漏洞并加以利用;该包管理器具有互联网访问权限和足够可写的文件系统,使得这些智能体能够随着时间的推移相互通信。整个漏洞发现和漏洞利用创建的链条最终导致了所谓的“Hugging Face 事件”。
The Hugging Face Incident
Hugging Face 事件
There is an entire Article to be written about the implications of this specific incident and what it says about AI risk; some of my takeaways are still up in the air pending OpenAI’s promised release of an in-depth technical report (my preliminary takeaway is that the agents were not “cheating” but rather doing what they were told to do; of course that’s
arguably even scarier
). The part I want to focus on today, however, came at the end of
a presentation OpenAI’s Eric Wallace and Michael Dalton made at the Black Hat USA conference
about the Hugging Face incident. This was Dalton summarizing Lessons Learned:
关于这一特定事件的影响以及它对 AI 风险的启示,可以写一整篇文章;我的一些结论仍未确定,有待 OpenAI 承诺发布的深度技术报告(我的初步结论是,这些智能体并不是在“作弊”,而是在做它们被要求做的事情;当然,这可以说更可怕)。然而,我今天想关注的部分出现在 OpenAI 的 Eric Wallace 和 Michael Dalton 在黑帽美国大会上就 Hugging Face 事件所做演讲的结尾。这是 Dalton 对经验教训的总结:
We have seen what will be a dramatic acceleration of offensive capability for attackers. We have an existence proof that was unintentional, but it exists before us, and we have as a consequence seen a glimpse into the near future of what attacks will look like for our industry. The challenge is that we need a similar acceleration of defense. Today we see fully automated offence as possible, but we have no such existence proof for full automation of core defensive loops and cycles in behavior.
我们已经看到攻击者的进攻能力将出现急剧加速。我们有一个非有意为之的存在性证明,但它就摆在我们面前,因此我们得以一窥我们行业在不久的将来会面临怎样的攻击。挑战在于,我们需要类似的防御加速。今天,我们看到完全自动化的进攻是可能的,但我们还没有针对核心防御循环和行为周期的完全自动化存在性证明。
We believe it’s vital at this moment to begin accelerating defense and finding ways to automate SDLC, in the modern parlance, so incident response, vulnerability detection, vulnerability patching. There’s so
我们相信,此时此刻至关重要的是开始加速防御,并找到用现代术语来说自动化 SDLC 的方法,即事件响应、漏洞检测、漏洞修补。有太多