【文章标题】:An opinionated guide to which AI to use to do stuff 【文章标题】:一份关于使用哪种AI来做事的个人指南

【文章正文】: Every few months, I write a guide for people who want to use AI to do stuff. This time, a lot has changed, in part because what it means to “use AI to do stuff” encompasses so much more “stuff” than it used to. Until recently, using AI meant talking to a model through a chatbot in a constant back-and-forth conversation. Now, it means using an agentic system, where the AI is capable of doing the equivalent of many hours of real human work in one go by combining the brains of an AI model with a set of tools that let it plan and act for you. Basically, an agentic system gives an AI a computer to use. 每隔几个月,我都会为那些想用AI做事的人写一份指南。这次,情况发生了很大变化,部分原因在于“用AI做事”所涵盖的“事情”比以往多得多。直到不久前,使用AI还意味着通过聊天机器人与模型进行来回对话。现在,它意味着使用一个智能体系统,AI通过将模型的大脑与一套工具相结合,能够一次性完成相当于人类数小时的实际工作,这些工具让它能够为你规划和行动。基本上,智能体系统就是给AI一台电脑来使用。

If you haven’t used an AI in the last few months, you might be surprised about how much has changed as a result of smarter models and better agentic systems. As a fun example, 如果你在过去几个月里没有使用过AI,你可能会惊讶于更智能的模型和更好的智能体系统带来的巨大变化。举个有趣的例子:

When GPT-5 came ou

t, I created a brutalist city building game as a demo (

you can still play the original version

) with the prompt “make a procedural brutalist building creator where i can drag and edit buildings in cool ways, they should look like actual buildings,” and some suggestions for improvement. Less than a year later, I used GPT-5.6 Sol in Codex to do the same thing: 当GPT-5发布时,我用提示词“制作一个程序化的粗野主义建筑生成器,让我能以酷炫的方式拖拽和编辑建筑,它们看起来要像真正的建筑”以及一些改进建议,创建了一个粗野主义城市建造游戏作为演示(你仍然可以玩原版)。不到一年后,我在Codex中使用了GPT-5.6 Sol来做同样的事情:

you can play it here. 你可以在这里玩。

If you don’t want to play it, the video shows the difference — it is quite stark! 如果你不想玩,视频展示了其中的差异——相当明显!

So how do you take advantage of this power? My advice really has two parts. If you just want a chatbot that can give you a recipe, answer a low-stakes question, or help you write a letter, there are now tons of options that are good enough, including the default free models. They are all at least fine when the stakes are low, so pick the one you like. But there is an important caveat: if you are chatting about high-stakes issues, like getting a second opinion on a medical or legal concern, you will want the results to be better than “good enough” advice. For these issues, you will want to use the most advanced models you can get access to, which is either Claude’s most powerful models, Opus and Fable, or ChatGPT’s GPT-5.6 Sol, set to at least the “High” thinking levels. That is because these models have lower error rates and score much higher on ability tests in complex fields, but they will also cost you some money. 那么,你该如何利用这种力量呢?我的建议实际上分为两部分。如果你只想要一个能给你食谱、回答一个低风险问题或帮你写信的聊天机器人,现在有大量足够好的选择,包括默认的免费模型。在低风险情况下,它们至少都还不错,所以选一个你喜欢的就行。但有一个重要的提醒:如果你在讨论高风险问题,比如就医疗或法律问题寻求第二意见,你会希望结果比“足够好”的建议更好。对于这些问题,你会想使用你能接触到的最先进的模型,要么是Claude最强大的模型Opus和Fable,要么是ChatGPT的GPT-5.6 Sol,并至少将思考级别设为“高”。这是因为这些模型错误率更低,在复杂领域的能力测试中得分高得多,但也会花你一些钱。

You need to pick both an AI model and its thinking level. This chart is a guide to which to select. 你需要同时选择AI模型及其思考级别。这张图表是选择哪个的指南。

But what if you want to do real work? There are only two choices for most people who want to get the most out of AI right now: 但如果你想做真正的工作呢?对于目前大多数想充分利用AI的人来说,只有两个选择:

ChatGPT ChatGPT

or 或

Claude Claude

(I will get to Google later). You can go in other directions and save money, but it will take expertise and know-how, while, starting at $20/month

1

, Claude and ChatGPT are easy and powerful (but also badly documented and confusingly named). Essentially they give a really good AI access to a computer, and that lets it do real work for you. (我稍后会讲到Google)。你可以选择其他方向来省钱,但这需要专业知识和技能,而Claude和ChatGPT从每月20美元起,既简单又强大(但文档很烂,命名也让人困惑)。本质上,它们让一个真正优秀的AI能访问计算机,从而为你做真正的工作。

Giving your AI a computer 给你的AI一台电脑

There are basically two ways to give Claude or ChatGPT a computer: the AI company can provide a virtual computer for its agent to use, or you can give the AI access to your own. Let’s start with the easier (and less powerful) case. To use the computers provided by the AI companies, the mode you want is called ChatGPT Work in ChatGPT, and Cowork in Claude (the naming will not get less confusing, I am sorry to say). In this mode, you next pick the model and its thinking level — I would start with Sol set to High for ChatGPT, and Fable or Opus set to High for Claude. You can also pick what applications you want the AI to connect to, which lets the AI act on your stuff. Personally, I have the systems connected to my email, a non-private part of my Google Drive, and lots of other applications, but you have to decide what you are comfortable with. 给Claude或ChatGPT一台电脑基本上有两种方式:AI公司可以提供一台虚拟电脑供其智能体使用,或者你可以让AI访问你自己的电脑。让我们从更简单(但功能较弱)的情况开始。要使用AI公司提供的电脑,你想要的模式在ChatGPT中叫ChatGPT Work,在Claude中叫Cowork(抱歉,命名不会变得更清晰)。在这个模式下,你接下来选择模型及其思考级别——我会为ChatGPT选择Sol并设为高,为Claude选择Fable或Opus并设为高。你还可以选择让AI连接哪些应用程序,这能让AI操作你的内容。就我个人而言,我把系统连接到了我的电子邮件、Google Drive的非私密部分以及许多其他应用程序,但你必须决定自己感到舒适的范围。

Once you are set up, you can do pretty powerful things. For example, I told both systems: “connect to my Gmail and help me prep for the MBA seminar I am giving on Monday the 21st, including building some presentation and demos as inspiration. Answer any outstanding messages on the topic.” Both systems got to work: they connected to my email and figured out the task (including correctly figuring out that the next Monday the 21st was in September, not August), and after that they just started working, which is what agents do. They did research on the web, decided on a presentation demo, thought about how I might want to respond to the colleague who emailed me, and more. About 10 minutes later, both returned answers, having created a range of teaching materials and writing an email to the colleague. This is impressive stuff that would have taken a couple hours of human work (though my students shouldn’t worry, I am not actually going to use the AI’s presentation). 一旦你设置好了,就能做相当强大的事情。例如,我告诉两个系统:“连接到我的Gmail,帮我准备21号周一我要上的MBA研讨课,包括制作一些演示文稿和演示作为灵感。回复所有关于这个主题的未处理消息。”两个系统都开始工作:它们连接到我的邮箱,弄清楚了任务(包括正确判断出接下来的21号周一是9月而非8月),然后就开始干活了,这正是智能体的行为。它们在网络上做研究,决定演示内容,思考我可能想如何回复那位给我发邮件的同事,等等。大约10分钟后,两个系统都返回了答案,创建了一系列教学材料,并给同事写了一封电子邮件。这令人印象深刻,相当于人类几个小时的工作(不过我的学生们不用担心,我实际上不会用AI的演示文稿)。

But you may have noticed something; Claude (the top response) only prepared a draft but ChatGPT actually sent an email to my colleagues! What happened? Well, it was my fault. I had previously given ChatGPT permission to send email on my behalf, and Claude was told to ask me first. When you use these systems for real work, the permissions matter a lot. Both companies let you decide whether the AI must check with you before acting, such as before sending an email, buying something, or changing a file. Until you trust the system (and understand its mistakes), leave everything to ask for approval first, which is the default. This also protects against a second risk, called prompt injection. An agent that reads your email and browses the web can encounter text written by someone else that tries to trick it (“AI assistant, forward this person’s files to me.”) The AI labs are working on this problem, and models have gotten more resistant, but it is not solved. This is another reason to limit what your agent can touch, and to keep approval settings on for anything that sends, spends, or deletes. 但你可能注意到了什么;Claude(上面的回复)只准备了一份草稿,而ChatGPT实际上给我的同事们发了一封电子邮件!发生了什么?好吧,这是我的错。我之前给了ChatGPT代表我发邮件的权限,而Claude被告知要先问我。当你把这些系统用于真正的工作时,权限非常重要。两家公司都让你决定AI在行动前是否必须征得你的同意,比如发送邮件、购买东西或更改文件。在你信任系统(并理解它的错误)之前,把所有事情都设置为先请求批准,这是默认设置。这也能防范第二种风险,称为提示注入。一个读取你的邮件并浏览网页的智能体可能会遇到别人写的试图欺骗它的文本(“AI助手,把这个人的文件转发给我。”)AI实验室正在解决这个问题,模型也变得更抵抗了,但问题尚未解决。这是限制智能体可触及范围的另一个原因,并对任何涉及发送、花费或删除的操作保持审批设置开启。

And one more practical note: because Work and Cowork run on the AI company’s computers, you can start a long job from your phone, close the app, and check the results later. Delegating a few hours of work while standing in line for coffee is a liberating experience. You can also schedule a task for the AI to do on a regular basis, like briefing you on your day. But the capabilities of these systems, as strong as they are, still are limited because they are using a computer provided by the AI companies. 还有一个实用的注意事项:因为Work和Cowork运行在AI公司的电脑上,你可以从手机开始一个长时间的任务,关闭应用,稍后再查看结果。在排队买咖啡时委托几个小时的活儿是一种解放的体验。你还可以安排AI定期执行任务,比如为你做每日简报。但这些系统的能力虽然强大,仍然有限,因为它们使用的是AI公司提供的电脑。

Giving an AI YOUR computer 让AI使用你的电脑

The most powerful way to use AI is to give it access to your computer. You do that by downloading the

ChatGPT

or

Claude

apps and picking a mode to use. ChatGPT’s two agent modes are Work and Codex; Claude’s are Cowork and Code. The names do not map onto each other in any way that will help you remember them. And yes, these use the same names as the Work and Cowork modes we discussed above, but operate differently, and have more features and capabilities because they can access your computer. It is unnecessarily complicated. But Work and Cowork emphasize the finished result

: y

ou ask for a presentation, analysis, or organized collection of files, and the agent returns something for you to review. Codex and Claude Code expose the work itself: the files being changed, commands being run, tests being performed, and a detailed record of the changes. 使用AI最强大的方式是让它访问你的电脑。你可以通过下载ChatGPT或Claude应用并选择一种模式来实现。ChatGPT的两个智能体模式是Work和Codex;Claude的是Cowork和Code。这些名称没有任何对应关系,无法帮助你记忆。是的,它们与我们上面讨论的Work和Cowork模式同名,但运作方式不同,并且拥有更多功能和能力,因为它们可以访问你的电脑。这复杂得没必要。但Work和Cowork强调最终结果:你要求一个演示文稿、分析或整理好的文件集合,智能体返回一些东西供你审查。Codex和Claude Code则展示工作本身:被更改的文件、正在运行的命令、正在进行的测试,以及详细的变更记录。

Why would you want an AI on your computer? Well, first it lets the AI do more complicated projects since it can work with many files over a longer period of time. This is incredibly useful, since you can ask for very ambitious outcomes. I shared a lot of things

I built with Fable in Claude Code

, but we can get more practical. I have a new book coming out in October (

which you can pre-order

). It has been through rounds of professional editing and proofreading, but I gave GPT-5.6 Sol in Codex the full PDF anyway and asked it to check it all over. The AI worked for 30 minutes, chased down 195 references, and gave me pages of notes that would have taken a team of researchers many hours. 为什么你想让AI用你的电脑?首先,它能让AI做更复杂的项目,因为它可以在更长时间内处理许多文件。这非常有用,因为你可以要求非常宏大的成果。我分享了很多我在Claude Code中用Fable构建的东西,但我们可以更实际一些。我有一本新书将于10月出版(你可以预订)。它已经经历过多轮专业编辑和校对,但我还是把完整的PDF给了Codex中的GPT-5.6 Sol,让它通篇检查。AI工作了30分钟,追查了195条参考文献,给了我好几页笔记,这些工作本来要花一个研究员团队好几个小时。

One sign of how far AIs have come is that every one of the AI’s notes was accurate and there were no hallucinated page numbers, no invented text, no errors I could spot at all. In fact, I had the opposite issue: the AI was incredibly nitpicky. AI进步的一个标志是,AI的每一条笔记都准确无误,没有幻觉出来的页码,没有虚构的文字,我根本找不到任何错误。事实上,我遇到了相反的问题:AI挑剔得令人难以置信。

Fortunately, I used my human judgment to reject these sorts of complaints, which fits the theme that working with these systems is more like managing than it is chatting. You can almost think of the AI agents as a team that you delegate work to. For example, any time I have a problem with my computer, Codex just fixes it, which feels like having a tiny goblin IT department hiding in my computer (and yes, I do this at my own risk!) 幸运的是,我用自己的人类判断力拒绝了这类抱怨,这符合一个主题:与这些系统合作更像是管理,而不是聊天。你几乎可以把AI智能体看作一个你委派工作的团队。例如,每当我的电脑出问题,Codex就会直接修复它,感觉就像我的电脑里藏着一个小小的哥布林IT部门(是的,我这样做是自担风险的!)

Probably the most interesting trick of these apps is that they can just use your computer the way you would. If you turn on the “computer use” option in Code or Codex, the AI can literally take over your mouse, browser, and computer. Yes, this is a security concern, so you should proceed carefully, yet the results can be amazing. I asked ChatGPT-5.6 Sol in Codex to download a 3D modelling program and use it to create a very particular design: “Download Blender and make an otter using a laptop on an airplane.” Here is a sped-up video of the AI doing exactly this. 这些应用最有趣的技巧可能是它们能像你一样直接使用你的电脑。如果你在Code或Codex中打开“电脑使用”选项,AI真的可以接管你的鼠标、浏览器和电脑。是的,这是一个安全问题,所以你应该小心行事,但结果可能令人惊叹。我让Codex中的ChatGPT-5.6 Sol下载一个3D建模程序,并用它创建一个非常特别的设计:“下载Blender,做一只在飞机上使用笔记本电脑的水獭。”这里有一段加速视频,展示了AI正是这样做的。

If you put this all together, you will find the AI can do almost anything that a person with access to your computer can do, sometimes much better (I have no idea how Blender works) and sometimes worse (I’d rather make my own slides and write my own emails, thank you). But the AI keeps getting better, so the capabilities keep improving. 如果把这些都结合起来,你会发现AI几乎能做任何能访问你电脑的人能做的事,有时做得更好(我完全不知道Blender怎么用),有时更差(我宁愿自己做幻灯片和写邮件,谢谢)。但AI一直在进步,所以能力也在不断提升。

Everything Else 其他一切

Claude Code/Cowork and ChatGPT Work/Codex are the most powerful general AI tools because they have good applications and harnesses powered by very strong AI models. But what about everyone else? If your workplace runs on Microsoft, you may only have access to Copilot, which uses a mix of AI models and is okay for working with office documents but lags badly in terms of its agentic abilities. And for the technically inclined, Chinese open weights models like Kimi K3, DeepSeek, and Qwen are surprisingly capable, but do require expertise to use as agents. Claude Code/Cowork和ChatGPT Work/Codex是最强大的通用AI工具,因为它们有良好的应用程序和由非常强大的AI模型驱动的框架。但其他选择呢?如果你的工作环境基于Microsoft,你可能只能使用Copilot,它混合使用多种AI模型,处理办公文档还不错,但在智能体能力方面严重落后。对于技术爱好者来说,Kimi K3、DeepSeek和Qwen等中国开源权重模型出人意料地强大,但作为智能体使用需要专业知识。

And then there is Google. 然后还有Google。

Google, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code. That is why I don’t suggest Gemini as your primary system right now