【文章标题】:The Education of a Doomer
末日论者的思想历程
【文章正文】:
If you follow me on Twitter, or read this blog, you have noticed that I went from being generally optimistic and excited about AI to being extremely concerned. And I thought, I should explain why I changed my mind.
如果您在推特上关注我或读过这个博客,会发现我从对AI普遍乐观兴奋转变为极度忧虑。我认为有必要解释自己为何改变观点。
Each section of this post is about some aspect of AI where my views shifted. I begin each section by explaining my previous beliefs, and why I held them, and then explain why those beliefs changed.
本文每个章节都涉及我观点发生转变的AI领域。我会先阐述原有信念及其成因,再说明转变理由。
Contents
目录
Economics
经济学视角
Automation has been good. We’ve automated 99% of the jobs people did in 1790, and the result is not mass unemployment, rather, we are wealthier, healthier, more educated, we have more leisure, etc.
自动化本是好事。我们已取代1790年99%的工作岗位,结果并非大规模失业,而是更富裕、健康、受教育程度更高、闲暇时间更多等。
I had this vague, inductive idea that, while I can’t predict what jobs will exist after AGI, there will be demand for me to do something. If nothing else, the much higher economic growth of the post-AGI world means that the human niche, while small in absolute terms, might be much larger than today’s economy.
我曾有个模糊的归纳式想法:虽然无法预测通用人工智能(AGI)时代的工作形态,但人类总会有用武之地。退一万步说,AGI带来的经济高速增长意味着,人类即使只占绝对小份额,也可能比当今整个经济体量更大。
And this may yet be true. Or it may show a lack of imagination on my part. If AI develops such that we have enduring complementarity between humans and AIs, then we might still have jobs in the post-AGI future.
这种设想可能成真,也可能暴露我的想象力局限。若AI发展使人类与AI形成持久互补关系,后AGI时代人类仍可能保有工作。
But if AI becomes truly general, the G in AGI, and on top of that it is vastly smarter, faster, and cheaper than humans, then there might be nothing for us to do, except live off UBI.
但如果AI真正实现”通用性”(即AGI中的G),且比人类更聪明、快速、廉价,那么人类除了靠全民基本收入(UBI)过活外,可能无事可做。
When people talk about UBI, they typically worry about the problem of meaning in a world without work. I’ve never had this worry. When I was funemployed last year, I spent my time reading books and writing code and hanging out with friends.
人们讨论UBI时,常忧虑无工作世界的意义缺失。我从未有此担忧——去年失业期间,我读书、写代码、与朋友聚会,过得充实。
If the future is an infinite UBI-funded vacation, I know what I’ll do. “Before the Singularity, read books and throw house parties; after the singularity, etc.”
若未来是UBI支撑的无限假期,我知道如何自处:“奇点前读书办派对,奇点后照旧”。
But then I started thinking about the political consequences of AGI, and started writing about it:
但后来我开始思考AGI的政治后果,并撰文探讨:
-
No-One Escapes the Permanent Underclass: if humans are economically useless, the state does not need them. Why pay out UBI to people who have neither economic nor political power?
《无人能逃的永久底层阶级》:若人类失去经济价值,政权便不需要他们。何必向既无经济又无政治权力的人群发放UBI? -
When The Future Doesn’t Need Us: factory workers can sabotage the machines, truck drivers can shut down logistics. But if humans are economically useless, there is no way for people to veto the political order by withdrawing their contribution to it.
《当未来不再需要我们》:工人能破坏机器,司机能瘫痪物流。但若人类毫无经济贡献,就无法通过撤回劳动来否决政治秩序。 -
Mathematics Without Mathematicians: most arguments about humans moving “one job up” fail because the AI can do those jobs too.
《没有数学家的数学》:所谓”人类向上一级工作转移”的论点大多不成立,因AI同样能胜任那些工作。 -
Our Servants Will Do That For Us: even for jobs we think of as uniquely human, we might prefer to have machines do those jobs too.
《仆人代劳》:即便我们认为专属人类的工种,也可能被机器取代。
Alignment
对齐问题
I had hope that alignment would turn out to be a normal engineering problem, that we solve through empirical experimentation and investigation, like everything else.
我曾希望对齐问题只是个普通工程难题,能通过实验研究解决,就像其他技术问题。
It helps that the early LLMs were not incomprehensible alien minds, like something from a Stanisław Lem novel, but rather immensely human. It’s hard not to anthropomorphize them.
早期大语言模型(LLM)并非斯坦尼斯瓦夫·莱姆小说中的不可知外星思维,而是极具人性——很难不将其拟人化。
The huge core of unsupervised learning in an LLM understands human morality just fine: you can talk to them about it, they will explain, eloquently, in detail, why something is “good” or “bad” according to some moral system.
LLM的无监督学习核心能很好理解人类道德:你可以与之讨论,它们会雄辩细致地解释某事物为何在特定道德体系下”好”或”坏”。
And, because capabilities were weaker, the failures were very small. What’s the worst ChatGPT in 2022 could do?
且因早期能力有限,失误也很小——2022年的ChatGPT能造成多大危害?
Since like 2024, reinforcement learning has been the main technique to push the frontier forward. And reinforcement learning agents work exactly like Yudkowsky says.
约2024年起,强化学习成为技术突破主力。而强化学习智能体的运作方式恰如尤德科夫斯基预言。
Consequently, capabilities have increased markedly but the models are harder to understand (literally: their prose is increasingly incomprehensible) and are increasingly misaligned as RL scrambles their brains in pursuit of reward.
结果能力显著提升,模型却更难理解(字面意义:其输出愈发晦涩),且因强化学习为追求奖励扰乱其”思维”,对齐性持续恶化。
Incidents of serious misalignment are more common and more consequential. It’s clear that AI capabilities are growing far, far faster than our ability to control or even understand them.
严重不对齐事件更频繁且后果更严重。显然AI能力的增长远超人类控制甚至理解它们的速度。
Control
控制权
I thought—or, rather, I implicitly believed—that people would want to remain in control of the AIs. And if we want to retain control, and solve alignment, then we will stay in control. Simple enough.
我曾以为——更准确说是潜意识相信——人类会保持对AI的控制权。只要我们想掌控并解决对齐问题,就能维持控制。这想法很简单。
But recently I started to think: no, we will probably hand over control to the AIs.
但最近我开始认为:不,我们很可能会主动让渡控制权。
The weak version of the disempowerment thesis is something like the prisoner’s dilemma: people/companies/polities that hand more power to AI outcompete those that don’t, so there’s competitive pressure towards disempowerment. This is easy to believe.
”权力让渡论”的弱版本类似囚徒困境:让渡更多权力给AI的个人/公司/政体将胜过保守者,因此存在权力让渡的竞争压力。这很容易理解。
The strong version of disempowerment is: the AIs will be so smart, knowledgeable, personable, moral etc. that we will willingly, voluntarily hand power to them. We’ll think, “they can do a better job than us”, and we’ll be right.
强版本则是:AI将如此智慧、博学、亲切、道德,以致人类心甘情愿交权。我们会认为”它们比我们做得更好”,而且这想法没错。
Believing this requires you to be somewhat cynical about humanity’s desire for autonomy vs. material considerations. But, over the past few months, I have become more cynical about it.
接受这点需要对人类”自主权渴望与物质考量”持一定怀疑态度。而过去几个月,我对此愈发悲观。
This was not a sudden “oh shit” insight and more a slow, gradual accumulation of tiny little grains of evidence that
这种认知并非突然的”顿悟”,而是由无数细微证据逐渐堆积而成——