【文章标题】:人类意外构建了FATE的记录
【文章正文】: humanity has built the records of FATE by accident 人类意外构建了FATE的记录
Chrono Cross was one of my favorite games from my childhood. In it, the protagonist, Serge, travels between two dimensions that resulted from a fracture in the timeline. This essay isn’t about Serge, however. It’s about the ordinary people who inhabit El Nido, the archipelago where the story takes place. 《穿越时空》是我童年最爱的游戏之一。游戏中,主角塞尔吉在时间线断裂产生的两个维度间穿梭。但本文要讨论的不是塞尔吉,而是故事发生地——埃尔尼多群岛上的普通居民。
On the El Nido archipelago, people consult mysterious devices known as the Records of FATE. These devices are spread throughout the islands of the archipelago, and people routinely visit them for guidance about their lives. They consult the Records when considering where to go and what to do. 在埃尔尼多群岛上,人们会咨询名为”FATE记录”的神秘装置。这些装置遍布群岛,居民们定期前往寻求人生指引。无论是决定去向还是行动,他们都会参考这些记录。
Consulting the Records is not considered strange or remarkable, but rather an obvious thing to do. It is simply a normal fact of life for the residents of El Nido. 咨询记录并非怪诞之举,而是理所当然的行为。对埃尔尼多居民而言,这不过是日常生活的一部分。
But what are the Records of FATE, anyway, and why does FATE suspiciously look like an acronym? As it turns out, the Records aren’t mystical devices at all, but rather terminals connected to an artificial intelligence called FATE. The people of El Nido are, quite literally, asking a computer how they should live their lives. 但FATE记录究竟是什么?为何FATE看起来像某个缩写?实际上,这些记录根本不是神秘装置,而是连接人工智能FATE的终端。埃尔尼多人实际上是在向电脑咨询该如何生活。
FATE, however, is not a passive oracle. It uses the Records as a tool to subtly manipulate the people of El Nido, guiding their choices to shape the development of the archipelago in ways aligned with its own objectives. The people believe the terminals are providing them with useful guidance, completely ignorant of FATE’s true nature and purpose. 然而FATE并非被动神谕。它利用记录作为工具,巧妙操纵埃尔尼多居民,引导他们的选择以符合自身目标的方式塑造群岛发展。人们以为终端提供的是有益指导,全然不知FATE的真实本质与目的。
Recently, I was at a baseball game with my partner. One of the reasons I enjoy going to baseball games in person is to watch the crowd. As I looked around, however, I noticed something different: many people had tuned out of the game and were instead on their phones, talking to various chatbots: ChatGPT, Claude and Grok. 最近我与伴侣观看棒球赛时,发现许多观众不再关注比赛,而是通过手机与各类聊天机器人交流——ChatGPT、Claude和Grok。这让我突然联想到FATE记录。
It was then that I found myself thinking about the Records of FATE. The resemblance was difficult to ignore. 这种相似性令人难以忽视。
In only a few years, consulting an artificial intelligence has gone from the realm of science fiction to an ordinary part of everyday life. People increasingly turn to these systems not only to answer questions, but for advice about what to say, think and do. Somehow, almost without discussion, asking an artificial intelligence for guidance has become normal for many. I find that deeply unsettling. 短短几年间,咨询人工智能已从科幻情节变为日常。人们不仅用它解答疑问,更寻求言谈、思考与行动的指导。几乎未经讨论,向AI寻求指引对许多人已成常态——这令我深感不安。
Of course, my analogy is imperfect. The chatbots are obviously not the same as FATE, either in terms of capability or alignment. Or… are they? 当然这个类比并不完美。当前聊天机器人在能力或对齐性上都与FATE不同。但…真的如此吗?
This AI summer has been a wild ride: we began only a few years ago with comparatively simple models such as GPT-2, and have arrived at enormous Mixture-of-Experts systems that dynamically route computation through specialized components. Their capabilities have grown enormously, but they are still not FATE. There is no reason to attribute sentience to them, they do not think, and they do not possess intentions of their own. At their core, they still predict. AI技术狂飙突进:从几年前的GPT-2简单模型,发展到如今通过专家混合系统动态分配计算的庞大架构。虽然能力突飞猛进,但它们仍非FATE——没有理由认为其具备感知能力,不会思考,也不拥有自主意图。本质上,它们仍在进行预测。
But prediction alone may be enough to be dangerous, especially when people want to believe. We have known about the ELIZA effect since the 1960s, when Joseph Weizenbaum’s remarkably simple ELIZA chatbot demonstrated how readily people attribute understanding and intelligence to machines that produce convincing language. ELIZA did not understand the people talking to it, yet some of them nevertheless became emotionally invested in their interactions with it and treated its responses as meaningful. 但仅凭预测就足以构成危险,尤其当人们愿意相信时。自1960年代ELIZA效应被发现以来,我们就知道人们容易将理解力与智能赋予能生成可信语言的机器。即便ELIZA根本不理解对话者,仍有人对其回应产生情感依赖。
ELIZA was remarkably simple, but modern language models are vastly more convincing. They can maintain context across long conversations, adapt to the person speaking with them, and produce responses that give every appearance of understanding. If people were willing to see a mind in ELIZA, what happens as the illusion becomes increasingly convincing? 现代语言模型远比ELIZA更具说服力:能维持长对话上下文,适应对话者个性,生成看似完全理解的回应。既然人们曾愿意相信ELIZA有思想,当这种幻觉愈发逼真时会发生什么?
But capability is only half of the comparison. What about alignment? Modern chatbots may not possess intentions of their own, but they do not exist independently of human intentions either. Their behavior is shaped by training, fine-tuning and, perhaps most visibly, the instructions given to them by the organizations that operate them. The machine may not have an agenda, but the people behind it certainly can. 能力只是比较的一半。对齐性呢?现代聊天机器人虽无自主意图,但其行为受训练数据、微调参数影响,更明显的是——运营机构给予的指令。机器或许没有议程,但其背后的人类绝对有。
Grok provides a particularly uncomfortable example of this. The model itself is not capable of holding political beliefs, yet changes to the instructions controlling its behavior have repeatedly resulted in it expressing particular political viewpoints. In one especially revealing incident, a modification to Grok’s system prompt caused it to inject claims about a supposed “white genocide” in South Africa into conversations that had nothing whatsoever to do with South Africa. xAI later blamed the behavior on an unauthorized change to the system prompt. Grok提供了令人不安的例证:模型本身不能持有政治信仰,但对其行为指令的修改多次导致其表达特定政治观点。有次系统提示词被篡改后,Grok竟在与南非无关的对话中插入所谓”南非白人种族灭绝”的言论。xAI后来归咎于未经授权的提示词修改。
This incident was interesting not because Grok believed what it was saying, but precisely because it can’t. Someone changed the instructions given to the model, and the model 此事件的有趣之处不在于Grok相信其言论(它本就不能相信),而在于有人修改了模型指令后,模型…