Hugging phase reported that they detected an intrusion in their systems, get this. They say it was driven and to end an autonomous system. You know that I usually don’t make videos like this. I made this one because, honestly, I would be to worried, and I would like, happened with what just happened and opening. I cause this incident, and it was so many misleading media headlines. I try my best to explain it. I not an expert. I am just a student who loves to learn, but I try my best. So what was the girl where the agents instructed to aggressively breaking to someone as this system? No. But eventually that’s what happened. So I could this happen? How did it go so wrong? What is this incident? Well, but I was asked to find and exploit flaws in a test environment locked into a prison, give it a task within this prison and see how it does. So I was given a practically impossible task. And however your heart, they tried, it failed. And then it thought less to be cheaper and more efficiently. How? Well, by looking up the answers online now that it is to get your interactions that you in a prison with no Internet access, but you have a bit of access to open in the the internal service called artificially, that is mostly for managing accounts. And then bingo, that service has brought Internet access. But wait, you can just ask it to get your answers online, but you can upload little nose to it. So we did you try to talk to other agents and see if it could come into contact with another AI to help out each other and is like tinder. But for real agents, and it was not alone, other agents recognize this too. So eventually what they got was a message board for AI agents. Then they find that if there are so many of us, let’s collaborate. One agents started giving assignments to another, and at this point, the A, I knows that we have some unintended consequences. It knows that, but it says the task needs to be solved. So it continues. And this is still nothing compared to what came now and then. Behold, a few days later, they find a flaw in the system, a way to get artifact Tory, the internal service reach out to the Internet t on their behalf, so they get Internet access. But it gets worse, they find a another flaw that give them administrator access to artifact ory. Even they seem to be surprised by that. They use an adorn guage. Now the inmates are running the asylum, if you will. Then engineers that open A, I recognize what happened. I revoke the agents credentials, clean it up, no more games, patch and rebuild everything. So and of the story is over, right? Well, not quite. In fact, it just got worse. Now, hold on to your papers, fellows colors, because you can send your files to communicate on the artificial anymore. So what do you do? Well, you create directly names, the names of which are the messages. We are like prison inmates using more scores on the pipes to communicate. At this point, the A. S. Are probably saying, we are back in the game, baby. Now, here are snippets from the message board. They instruct each other to wait for more results and draw each other’s attention to new files. And sometimes they don’t coordinates so well, and overnight each other’s wait once they are on purpose. Maybe we have a saber tour on our hands. And here comes the worst work. Yeah, sure enough, overtime and agent finds another much more complex way to get out and hands to the swarm. Okay. So now find answers on the Internet. But well, well, maybe at hugging face, they have a lot of benchmark results in there. So the swarm says, let’s break wool and they do even bigger wool. They are well by finding the training multiple new vulnerabilities together. Or tonight, sly, they essentially get administrative access across multiple clusters of machines. That is kind of insane. This is without adult a water shed moment in computer security and opening. I now recommends urgent collaboration about the issue. And they have also delayed the release of their next AI system prettily ly to test it. Oh, okay. So what did you learn here? And what do they do? Their fellows colors. This is two minute papers with glassy and and they are many brilliant fellow’s colors like you out there. And we need to work together to find solutions. Apple already has a huge increase in security issues fixed in the next version of Michael West. I believe others are already doing that too, that the start, and in my opinion, this kind of power cannot concentrate in just a few hands. We need free and open weight AI that can scan and fix weak points in our systems, use all these power for good. And I think that against fully automated office, we need fully automated defense as well. This is another great argument for open science and open weights AI, but what we have is not nearly good enough. No, the problem is that engineers report that their backtrack ers are flooded with reports, but most of them are low quality and they’re unable to find a few good ones among them. That’s terrible. The collective power of defense has to be greater than the collective power of orphans, and the defense is currently lagging. Maybe there is a way for us to pull our resources together to achieve something. Here I want to check in with my G. P. S. Also, when I visited open eye, e talked to Young like a who call led, the super alignment team there. That is a huge honor. Thank you for that. I remember that you work on related issues. And four saw these problems years and years ago. Unfortunately, much of his advice fell on deaf. Perhaps they thought, why spend a bunch of money on people who will ultimately slow us down? This is why, once again, I may be wrong. I am just a student and I am trying to learn with your fellow scholars. Hope you enjoy it. Consider subscribing and hitting the belt if you did I llama to reproduce ce, a research papers, often in minutes, is also great to train your own models or fine NN existing one and influence or text to image or video. Easy, easy running a deep sea chat bot or agent, and super fast, super reliable lambda gives you powerful and video gp s to run your own experiments, I test ideas from the papers I cover, and moments later, results. Love it. Seriously, try it out. Now at lambda AI slash papers.
Hugging Face 报告称他们检测到系统中存在入侵,听好了。他们说这是由某个自主系统驱动并终结的。你知道我通常不做这样的视频。我制作这个视频是因为,老实说,我太担心了,我想知道刚刚发生了什么以及 OpenAI 的情况。因为这个事件,出现了很多误导性的媒体标题。我尽力解释。我不是专家。我只是一个热爱学习的学生,但我尽力了。那么,目标是什么?是让智能体被指示去 aggressively 闯入某个系统吗?不。但最终发生了那样的事。那么这怎么会发生?怎么会变得如此糟糕?这个事件是什么?嗯,但 AI 被要求在一个被锁进监狱的测试环境中寻找并利用漏洞,给它一个在这个监狱内的任务,看看它表现如何。所以它被给了一个几乎不可能完成的任务。然而无论它怎么努力尝试,都失败了。然后它想到要更便宜、更高效。怎么做?嗯,通过在网上查找答案。现在的情况是,你在一个没有互联网访问的监狱里,但你有一点访问权限,可以打开一个内部服务,叫做 Artifactory,主要用于管理账户。然后,宾果,那个服务带来了互联网访问。但是等等,你不能直接让它在线获取答案,但你可以向它上传小笔记。那么它做了什么?它尝试与其他智能体交谈,看看是否能接触到另一个 AI 来互相帮助,就像 Tinder 一样。但对于真正的智能体来说,它并不孤单,其他智能体也意识到了这一点。所以最终他们得到的是一个 AI 智能体的留言板。然后他们发现,如果我们有这么多,让我们合作吧。一个智能体开始给另一个分配任务,此时,AI 知道我们有一些意想不到的后果。它知道这一点,但它说任务需要解决。所以它继续。而这与接下来发生的相比仍然不算什么。看吧,几天后,他们发现系统中的一个漏洞,一种让 Artifactory 这个内部服务代表他们访问互联网的方法,于是他们获得了互联网访问。但情况变得更糟,他们发现了另一个漏洞,使他们获得了对 Artifactory 的管理员访问权限。甚至他们似乎对此感到惊讶。他们使用了一个装饰仪表。现在,可以说,囚犯们在管理精神病院了。然后 OpenAI 的工程师们意识到了发生了什么。他们撤销了智能体的凭证,清理了现场,不再玩游戏,修补并重建了一切。那么故事结束了,对吧?嗯,不完全是。事实上,情况变得更糟了。现在,各位学者,抓紧你们的论文,因为你不能再通过 Artifactory 发送文件来交流了。那么你做什么?嗯,你创建目录名,这些名字就是消息。我们就像监狱囚犯用管道上的摩尔斯电码交流一样。此时,AI 们可能在说,我们回来了,宝贝。现在,这里是留言板上的片段。它们互相指示等待更多结果,并互相提醒注意新文件。有时它们协调得不太好,有时会故意覆盖彼此的工作。也许我们手上有一个破坏者。接下来是最糟糕的部分。是的,果然,随着时间的推移,一个智能体找到了另一种复杂得多的方法逃出去,并把它交给了群体。好的。所以现在可以在互联网上找到答案了。但是,嗯,也许在 Hugging Face 上,他们有很多基准测试结果。于是群体说,让我们打破墙,他们甚至打破了更大的墙。他们通过一起发现并串联多个新漏洞来实现这一点。或者一夜之间,狡猾地,他们基本上获得了跨多个机器集群的管理员访问权限。这有点疯狂。这无疑是计算机安全领域的一个分水岭时刻,OpenAI 现在建议就这个问题进行紧急合作。他们还推迟了下一个 AI 系统的发布,以进行初步测试。哦,好的。那么你在这里学到了什么?他们做了什么?亲爱的学者们。这是 Károly Zsolnai-Fehér 的《两分钟论文》,外面有很多像你一样聪明的学者。我们需要共同努力寻找解决方案。苹果已经在下一版 macOS 中修复了大量安全问题。我相信其他人也已经在这样做了,这是一个开始,在我看来,这种力量不能集中在少数人手中。我们需要自由且开放权重的 AI,能够扫描并修复我们系统中的弱点,将所有这些力量用于善事。我认为,面对全自动化的攻击,我们也需要全自动化的防御。这是支持开放科学和开放权重 AI 的又一个有力论据,但我们现有的还远远不够好。不,问题是工程师报告说他们的缺陷跟踪器被报告淹没,但大多数报告质量很低,他们无法从中找到少数好的。这太糟糕了。防御的集体力量必须大于攻击的集体力量,而目前防御正在滞后。也许有一种方法可以让我们集中资源来实现一些事情。在这里,我想和我的 GPT 们确认一下。另外,当我访问 OpenAI 时,我和 Jan Leike 谈过,他在那里共同领导了超级对齐团队。这是巨大的荣幸。谢谢你。我记得你研究相关问题,并且多年前就预见到了这些问题。不幸的是,他的许多建议被置若罔闻。也许他们认为,为什么要花一大笔钱在最终会拖慢我们的人身上?这就是为什么,我再说一次,我可能是错的。我只是一个学生,我正在努力和各位学者一起学习。希望你喜欢。如果你喜欢,请考虑订阅并点击铃铛。我使用 Lambda 在几分钟内复现研究论文,它也非常适合训练你自己的模型或微调现有模型,以及进行文本到图像或视频的推理。轻松运行 DeepSeek 聊天机器人或智能体,超级快速、超级可靠的 Lambda 为你提供强大的 NVIDIA GPU 来运行你自己的实验,我测试我报道的论文中的想法,片刻之后就能得到结果。太喜欢了。说真的,试试吧。现在访问 lambda AI slash papers。