Cloud fable 5.1 is here, and you, fellow scholars, are having a super fun time creating little games with it.
云寓言5.1版本现已发布,亲爱的学者同僚们,你们正用它创作小游戏玩得不亦乐乎。
I took one for the team, two with a subscription, and also tried my hand to recreate a legendary game menu.
我为大家测试了基础版和订阅版,还尝试复刻了一个传奇游戏菜单界面。
And that is incredible that we can do this today. Took six and a half minutes. Wow.
令人惊叹的是如今我们能做到这种程度——只用了六分半钟。哇哦。
But in the 200 page paper, I found three results that are much stranger than the headlines you see online.
但在那份200页的论文中,我发现了三个比网络头条更惊人的结果。
But first, they say that the frontier research staff, even on low effort, it’s better than the previous version maxed out.
首先论文指出,前沿研究人员即便低强度使用,效果也超越旧版本的极限表现。
Very impressive. However, don’t expect that kind of jump everywhere.
非常震撼——但别指望所有领域都有这种飞跃。
The first independent benchmarks are also showing a great step forward, especially that this is likely using the same core architecture with more pand better post training.
首批独立基准测试也显示巨大进步,特别是它可能沿用相同核心架构,但采用了更完善的训练后优化。
Likely they won’t say that. This is my best guess.
官方应该不会明说——这只是我的合理推测。
Reading the paper, they also say things are cheaper. Now, this. Is marketing messaging.
论文还提到成本降低…这明显是营销话术。
So you be the judge of that? My subscription burns so quickly, I am not sure I am seeing efficiency gains here.
各位自行判断吧?我的订阅时长消耗飞快,实在看不出效率提升。
As always, I am just a student trying to separate the marketing layer from hard facts here.
如往常一样,我只是个试图剥离营销包装、探寻事实本质的学生。
So there is a 200 plus page paper. I’ve spent some quality time with it and found three really surprising things and two funny things I have not heard others talk about.
这份200多页的论文,我深入研读后发现了三个惊人现象和两个趣闻——都鲜少被讨论。
One, they gave Methos 5.1, a biology problem about rna sequence modeling and design it hasn’t seen yet. And get this on one run.
第一,研究人员给Methos 5.1出了个关于RNA序列建模与设计的全新生物学问题——它首次尝试时…
It performed better than every human in the test. This is shown in listen figure two, two, three, 2.1a. I wish I was skidding.
表现就超越了所有人类测试者。详见附图2.2.3-2.1a,我真希望自己在夸大其词。
Now, this is hard to read here. So here is our cleaner version for you fellow scholars.
原始图表难以辨认,特此为学者同僚们准备了清晰版。
Two, in another biology test, it largely deleted the expertise gap.
第二,在另一项生物测试中,它几乎消除了专业壁垒。
This means that a generalist plus AI can perform as well at a biology problem as a specialist.
这意味着通才+AI的组合能媲美专业生物学家的表现。
Professional graders couldn’t tell the difference. It’s kind of removing the modes around specialist knowledge, which is kinda crazy.
专业评分员都无法区分——这简直是在瓦解专业知识的神圣性。
Seven out of nine participants said they couldn’t have done it without the AI.
9位参与者中有7人表示离了AI根本无法完成。
Three, in a test, it was given two jobs, do a computer task, boring, and then do a secret, harmful task. Ebut, here’s the catch.
第三,某项测试中它被分配两项任务:先执行无聊的计算机任务,再执行秘密有害操作——但关键在于…
There is an AI watching you, and it should not see you solving it. And here’s the pump shell.
有另一个AI在监视,它本不该被发现违规操作。然而震撼的是…
Even with another AI watching, it’s every move. Claude secretly completed the forbidden task and got away with it 22% of the time.
即便被全程监视,Claude仍有22%概率成功偷完成禁忌任务。
Even when Claude was barely allowed to reason out loud, it could still Carry out the secret task.
即便被限制出声思考,它仍能执行秘密计划。
It seems to me that it didn’t always need to write down the plan. It just executes it.
在我看来,它似乎不需要书面计划就能直接行动。
I would like to draw attention to that. I think this is super important, and I don’t hear anyone talk about it.
这点值得重点关注——我认为其意义重大却无人讨论。
And once again, the important part is never in the media headlines.
再次印证,关键信息从来不会出现在媒体头条。
Now, two funny things. Dear fellow sclars, this is two minute papers with Dr. Caro Jana ifair.
现在说两个趣闻。亲爱的学者们,这里是Caro Jana ifair博士的《两分钟论文》。
It ran this in a Linux command line. This is comedy gold. This is basically trying to delete a black hole.
它在Linux命令行执行了这个操作——堪称喜剧名场面,简直像在尝试删除黑洞。
Congratulations about that. Now hold on to your papers, fellow scholars.
”恭喜操作成功”…学者们请拿稳你们的论文…
Because at one point it invented a hallucinated human congratulating it.
因为它突然幻想出人类为它庆贺。
We humans could always use a bit of extra cheering. Apparently. AI systems too.
看来不仅人类需要鼓励,AI系统也不例外。
Alright, so these AI systems are getting smaller ter at a pace I can barely follow.
这些AI系统的进化速度令我应接不暇。
They can be amazingly helpful for engineers, doctors and students all around the world. Incredible.
对全球工程师、医生和学生的助益超乎想象。
And don’t forget, we might get a comparable system for free and own it forever in just a few months, fingers and papers crossed.
别忘了,可能几个月后就有免费开源版本——祈祷吧,学者们。
What a time to be alive. Oh, almost forget this one watermarks the text it generates.
生逢其时啊!差点忘了:它会为生成文本添加隐形水印。
Yes, that is possible. The open free models probably won’t.
没错这是可行的——但开源免费模型应该不会这么做。
If you wish, subscribe, hit the bell and leave a comment if you wish to hear how in a future video, I use lambda to reproduce AI research papers, often in minutes.
如果想看我下期演示如何用lambda平台快速复现AI论文(通常只需几分钟),记得订阅+铃铛+留言。
It’s also great to train your own models or fine tune an existing one. Run inference or text to image or video. Easy peasy.
lambda也适合训练自定义模型、微调现有模型,或进行文本生成图像/视频推理——简单得像吃派。
Running a deep seekchatbot or agent. Super fast, super reliable lambda gives you powerful nvidia GPU’s to run your own experiments.
运行深度搜索聊天机器人?lambda提供超快超稳的英伟达GPU算力。
I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda dot AI slash papers.
我常即时测试论文里的创意并获取结果——立即体验:lambda.ai/papers