【文章标题】:The twilight of the chatbots
【文章标题】中文翻译:聊天机器人的黄昏
【文章正文】:
If you feel like things are accelerating in AI, you are probably right. Better AI models from the leading American AI labs have been releasing more quickly than ever (though government interventions stopped access temporarily to two of the most powerful models, Claude Fable and GPT-5.6).
如果你感觉人工智能领域的发展正在加速,你可能是对的。美国领先人工智能实验室发布更强大模型的频率比以往任何时候都要高(尽管政府干预措施曾暂时阻止了人们对两个最强大模型——Claude Fable和GPT-5.6——的访问)。
But it isn’t just release timing. The evidence points to accelerating capability gains as well (though the frontier stays jagged, and AIs remain weak in many places). This is especially obvious when we look at the ability of AIs to do real work. There are a few good assessments that try to measure how much human work AIs can do. Two of the most famous, from METR and the UK’s official government AI Security Institute, estimate the amount of human programmer hours’ worth of effort the AI can do with a single prompt. GDPval compares human experts in many fields to AI performance using professional judges. They are all increasing at a better than exponential rate.
但加速的不只是发布节奏。证据同样表明能力提升也在加速(尽管前沿水平依然参差不齐,人工智能在许多领域仍然薄弱)。当我们审视人工智能完成实际工作的能力时,这一点尤为明显。目前有一些不错的评估试图衡量人工智能能完成多少人类工作。其中最著名的两个评估分别来自METR和英国政府官方的人工智能安全研究所(AI Security Institute),它们估算人工智能仅凭一个提示词就能完成相当于多少人类程序员工时的工作。GDPval则利用专业评审将多个领域的人类专家与人工智能的表现进行对比。这些指标都在以超指数级的速度增长。
Another organization doing similar experiments, Epoch, recently found Opus 4.7, working on its own for 14 hours, was able to build a software package that would take 2-17 weeks of human engineering work (it cost $251 in tokens). Again, AI systems cannot pass every test, nor are they always cheap to run, but they are definitely improving at a very rapid rate. In my own experiments, I found Fable was able to work autonomously for 9 hours to execute on very complex software projects that would have taken a team well over a week to do.
另一个开展类似实验的机构Epoch最近发现,Opus 4.7在自主运行14小时后,能够构建一个需要人类工程师2至17周才能完成的软件包(其令牌成本为251美元)。同样,人工智能系统并非能通过所有测试,运行成本也并非总是低廉,但它们确实在以极快的速度进步。在我自己的实验中,我发现Fable能够自主工作9小时,执行非常复杂的软件项目,而这些项目若由人类团队完成,需要整整一周以上的时间。
So far, I have focused on the frontier models, those with the highest “intelligence.” They are made by three American companies — Anthropic, OpenAI, and Google (though it has been a while since Google has released a new model). But there is a second set of near-frontier AI models that typically lag 6-12 months behind the frontier, all of which are from China. These are open weights models, which means that anyone can use or modify them after release (as opposed to the frontier models which are proprietary). That makes them quite cheap to operate. They, too, are climbing up an exponential improvement curve, though lagging the American closed models. You can see this in my graph of AI performance in a test called AA-Briefcase, which simulates a complex multi-week consulting engagement where AI has to do many kinds of analysis. The open-weights Chinese models (other countries produce open weights models, but none are near the frontier) are on their own exponential curve, behind closed US models.
到目前为止,我关注的都是前沿模型,即那些拥有最高”智能”的模型。它们出自三家美国公司——Anthropic、OpenAI和谷歌(尽管谷歌已经很长时间没有发布新模型了)。但还有第二类接近前沿的人工智能模型,通常落后前沿水平6至12个月,它们全部来自中国。这些是开放权重模型,意味着发布后任何人都可以使用或修改(这与专有的前沿模型形成对比)。这使得它们的运行成本相当低廉。它们同样在沿着指数级改进曲线攀升,尽管落后于美国闭源模型。你可以在我绘制的AA-Briefcase测试的人工智能性能图表中看到这一点。该测试模拟一个复杂的、为期数周的咨询项目,人工智能需要完成多种类型的分析。中国的开放权重模型(其他国家也生产开放权重模型,但没有一个接近前沿水平)正沿着自己的指数曲线前进,落后于美国闭源模型。
But abstract graphs only get you so far, and they can hide how jagged the frontier is (and also the fact that the open weights models, while very impressive, do not always perform as well as their benchmarks would indicate). To get real insight, you need to try using AI for different use cases and rigorously assess how good they are in the areas that matter to you. As a fun example, I created a test where AIs have to build an interactive simulation of a harbor evolving over time.
但抽象的图表只能带你走到这一步,它们可能掩盖前沿水平的参差不齐(以及这样一个事实:开放权重模型虽然令人印象深刻,但实际表现并不总是像其基准测试所显示的那样出色)。要获得真正的洞见,你需要尝试在不同用例中使用人工智能,并严格评估它们在你所关心的领域中的表现。举个有趣的例子,我设计了一个测试,让人工智能构建一个港口随时间演变的交互式模拟。
You can play with all the result here.
你可以在这里体验所有结果。
I think it gives an interesting perspective on how much models can differ from each other in areas like design, stylistic approach, and even judgement. As systems do ever longer tasks, these hard-to-benchmark factors become more important.
我认为这提供了一个有趣的视角,让我们看到模型之间在设计、风格取向甚至判断力等方面的差异有多大。随着系统执行的任务越来越长,这些难以用基准衡量的因素变得更加重要。
The way we use AI is changing
我们使用人工智能的方式正在改变
As AIs can do longer and longer tasks, the way people are using AI is changing. Until recently, the dominant way to use AI was as a co-intelligence. You would ask the AI to do something, check the results, and then ask for it to do the next step of your job. By careful prompting and human attention, you could guide AIs to do complex and long-term tasks.
随着人工智能能够执行越来越长的任务,人们使用人工智能的方式也在改变。直到不久之前,使用人工智能的主要方式还是将其作为协作智能。你会让人工智能做某件事,检查结果,然后让它继续完成你工作的下一步。通过精心设计的提示词和人工关注,你可以引导人工智能完成复杂而长期的任务。
This approach to using AI is still common and useful, but, increasingly, it is not the way AI is being used for valuable work. Long-running, smart, and self-correcting AI systems do not need constant human intervention, and they require a different way of working (this is also the subject of my upcoming book, Co-Existence, which you might want to pre-order here). And, as opposed to chatbots, agents come with extra machinery: harnesses that give the AI access to tools and an environment to act in, and apps built for agents like Claude Code or OpenAI’s Codex. As a result, the already increasing ability of AI models can be improved still further by a good harness or app.
这种使用人工智能的方式仍然普遍且有用,但越来越多的情况下,它已不再是人工智能被用于高价值工作的方式。长时间运行、智能且能够自我纠错的人工智能系统不需要持续的人工干预,它们需要不同的工作方式(这也是我即将出版的新书《共存》(Co-Existence)的主题,你可以在此处预订)。此外,与聊天机器人不同,智能体附带额外的机制:为人工智能提供工具访问权限和行动环境的框架,以及为智能体构建的应用程序,如Claude Code或OpenAI的Codex。因此,人工智能模型本已不断提升的能力,还可以通过良好的框架或应用程序得到进一步改进。
So work is increasingly about assigning work to agents, rather than working together with chatbots. A joint study by OpenAI and academic economists shows how quickly this is happening inside their own organization. Critically, it isn’t just coders who are using agents. Legal, HR, and other non-tech functions have adopted agents at nearly the same rate. OpenAI may be a sort of canary in the coal mine for what will happen elsewhere in work.
因此,工作越来越多地变成将任务分配给智能体,而不是与聊天机器人协作。OpenAI与学术经济学家联合开展的一项研究显示,这一转变在其组织内部正在以多快的速度发生。关键在于,使用智能体的不仅仅是程序员。法务、人力资源和其他非技术职能部门采用智能体的速度几乎相同。OpenAI可能就像煤矿中的金丝雀,预示着一场工作领域的变革将在其他地方发生。
Increasingly, work at OpenAI looks like managing AI. A quarter of OpenAI workers have at least four agents running at one time every week. And, as coding is done by AIs in specialized harnesses and apps, other roles start to become coders of a sort. And they are good at it. A separate study of Claude Code users found that software engineers had a similar success rate to other professions when actually using Claude code on coding tasks.
OpenAI内部的工作日益呈现出管理人工智能的形态。四分之一的OpenAI员工每周至少同时运行四个智能体。而且,随着编码由人工智能在专用框架和应用程序中完成,其他岗位的角色也开始变成某种意义上的程序员。而且他们做得很好。另一项针对Claude Code用户的研究发现,在实际使用Claude Code处理编码任务时,软件工程师的成功率与其他职业相近。
What actually mattered was not the profession of the user, but their expertise. The more domain experience someone had, the more successful they were in using Claude Code in that domain. And, even more interestingly, the more useful output they got from Claude from each prompt.
真正重要的不是用户的职业,而是他们的专业能力。一个人在某个领域积累的经验越多,他在该领域使用Claude Code就越成功。更有趣的是,他们从Claude的每个提示词中获得的输出也越有价值。
We are moving from a world where non-experts use chatbots to fill in gaps to one in which experts use agents to get work done. And the best way to use agents is to think of yourself as a manager.
我们正在从一个非专家使用聊天机器人填补空白的时代,走向一个专家使用智能体完成工作的时代。而使用智能体的最佳方式,就是把自己想象成一位管理者。
A moment in time
一个特定的时刻
Being on an exponential means each change over a fixed window is larger than the one before it. If your organization wrote an AI plan any time before the winter of 2025, it described a system that could do a couple of hours of work with a fairly high error rate. A few months later, you can get sixteen hours or more of work from a single prompt. This is why AI keeps feeling like it is making leaps, even though it is a curve on a graph, we keep experiencing a steady doubling of capability as a series of shocks. We are very bad at feeling exponentials from the inside, and we are currently inside one.
处于指数曲线上意味着,在固定的时间窗口内,每一次变化都比前一次更大。如果你的组织在2025年冬天之前的任何时候制定了一份人工智能计划,那么该计划所描述的系统只能完成几个小时的工作,且错误率相当高。而仅仅几个月后,你用一个提示词就能获得十六小时甚至更多的工作量。这就是为什么人工智能总让人感觉在飞跃——尽管从图表上看它只是一条曲线,但我们却将能力的持续倍增体验为一次次冲击。我们非常不擅长从内部感知指数曲线,而目前我们正处于这样一条曲线之中。
I think this also explains the turbulence around AI better than the usual stories about hype. AI is not capable of being a real cybersecurity threat until suddenly it is, causing sudden and improvised policy changes at the highest level of government. Markets discount whether AI might threaten to undermine a business model until suddenly it can, leading to massive swings in stocks. These lurches these get read as signs of an immature field that will eventually settle into something stable. I don’t think it is going to settle anytime soon. The instability is what happens when institutions that move at the speed of people (or worse, committees) try to track a capability curve that is very much not human in nature. And as long as we are on some sort of exponential, and for as long as it lasts, the gap only widens.
我认为这比那些关于炒作的常见说法更能解释人工智能领域周围的动荡。人工智能起初不具备构成真正网络安全威胁的能力,直到突然间它具备了,从而在最顶层的政府中引发突然而临时的政策变化。市场起初不理会人工智能可能颠覆某个商业模式的威胁,直到突然间它确实做到了,导致股市剧烈波动。这些剧烈摇摆被解读为一个不成熟领域的标志,人们认为它最终会趋于稳定。但我不认为它会在短期内稳定下来。这种不稳定性恰恰是那些以人的速度(甚至更糟,以委员会的速度)运转的机构,试图追赶一条本质上非人类的能力曲线时必然发生的结果。只要我们处于某种指数曲线上,只要这种状态持续下去,鸿沟就只会越来越大。
Subscribe now
立即订阅
Share
分享