【文章标题】:Weekly Dose of Optimism #209
乐观主义周剂量第209期
【文章正文】:
Hi friends 👋,
朋友们好👋,
Happy Friday, Happy Labor Day Weekend to my fellow Americans, and welcome to our 209th Weekly Dose of Optimism.
周五快乐,祝我的美国同胞们劳动节周末愉快,欢迎来到第209期乐观主义周剂量。
Right on cue, we have a week so full of celebration-worthy labor from the good guys that I don’t even know where to fit it all. We have a couple extra stories above the fold and a whole heaping for not boring world subscribers below.
恰逢其时,本周充满了来自好人们的值得庆祝的劳动成果,我甚至不知道该如何全部容纳。我们有一些额外的故事放在显眼位置,还有一大堆内容留给不觉得世界无聊的订阅者。
If this is what acceleration feels like in the summer, hold onto your butts in the fall.
如果这就是夏天的加速感,那秋天可得抓紧了。
Let’s get to it.
我们开始吧。
This Week’s Weekly Dose is brought to you by…
本周的乐观主义周剂量由以下赞助商为您呈现……
Reducto
Reducto
The Only Document Parsing Product You’ll Need
您唯一需要的文档解析产品
Reducto
Reducto
is the document ingestion layer trusted by AI teams at Scale AI, Harvey, Toast, and Vanta to reliably turn complex files into structured, grounded data for production AI systems at scale. Processing more than a billion pages per month,
是Scale AI、Harvey、Toast和Vanta等AI团队信赖的文档处理层,能够可靠地将复杂文件转化为结构化、有依据的数据,用于大规模生产AI系统。每月处理超过10亿页文档,
Reducto
Reducto
has built its reputation on the long tail of difficult documents that break conventional OCR: dense tables, unusual layouts, handwriting, low-quality scans, strikethroughs, and formatting-dependent information.
凭借那些让传统OCR崩溃的困难文档的长尾建立了声誉:密集表格、非常规布局、手写内容、低质量扫描、删除线以及依赖格式的信息。
Its
它的
new frontier parsing model, r-1
新前沿解析模型r-1
, combines layout detection, reading order, tables, formatting, grounding, and granular citations in a single request. It reduces parsing errors by up to 20%, improves latency at high volumes, and costs 1¢ per page all-in.
将布局检测、阅读顺序、表格、格式、依据和细粒度引用结合在一个请求中。它将解析错误减少高达20%,在高负载下改善延迟,每页综合成本仅1美分。
Already using another parser?
已经在使用其他解析器?
Apply for up to $5,000 in migration credits
申请高达5000美元的迁移积分
and hands-on support to benchmark r-1 against your current stack on the documents that matter most.
以及实际操作支持,在最重要的文档上将r-1与您当前的堆栈进行基准测试。
Apply at reducto.ai/migrate
在reducto.ai/migrate申请
(Early Bonus)
(早期福利)
Any Human Ever
任何人类
h/t Gina Gorlin
感谢Gina Gorlin
I’ve been saying for a while that the best thing the pro-progress people could build is a Counterfactual Machine. Run a simulation in which you remove certain technologies from humanity’s toolkit and let people experience that world in VR.
我一直在说,支持进步的人们能建造的最好的东西是一台反事实机器。运行一个模拟,从人类的工具包中移除某些技术,让人们通过VR体验那个世界。
This isn’t quite that, but it’s pretty darn close. Someone built a site that lets you draw a life at random from any of the over 100 billion people who have ever lived. As they wrote, “most lives ever lived were Asian, poor, and brief
这虽然不是完全一样,但也非常接近了。有人建了一个网站,让你从曾经生活过的超过1000亿人中随机抽取一个生命。正如他们所写,“大多数曾经生活过的生命都是亚洲人、贫穷且短暂
,” and indeed, the life that I pulled was an Asian woman who lived in the 11th century and lost four of her six kids before age 5.
”,确实,我抽到的生命是一位生活在11世纪的亚洲女性,她在5岁前失去了六个孩子中的四个。
Try it for yourself
自己试试看
.
。
In a 2016 speech, President Barack Obama said, “If you had to choose one moment in history in which you could be born, and you didn’t know ahead of time who you were going to be — what nationality, what gender, what race, whether you’d be rich or poor, gay or straight, what faith you’d be born into — you wouldn’t choose 100 years ago. You wouldn’t choose the fifties, or the sixties, or the seventies. You’d choose right now.” He was drawing on John Rawls concept of the Veil of Ignorance (and if you haven’t, stop what you’re doing and read Scott Alexander’s
在2016年的一次演讲中,巴拉克·奥巴马总统说:“如果你必须选择历史上的一个时刻出生,而你又不知道你将成为谁——什么国籍、什么性别、什么种族、富或穷、同性恋或异性恋、什么信仰——你不会选择100年前。你不会选择五十年代、六十年代或七十年代。你会选择现在。”他借鉴了约翰·罗尔斯的“无知之幕”概念(如果你还没读过,停下你手头的事,去读斯科特·亚历山大的
Being John Rawls
《成为约翰·罗尔斯》
).
)。
People love to talk about the challenges of modern life, and there certainly are some, but there’s also never been a better time to be alive in human history. The only better times, I suspect, all lie in the future, thanks to the types of stories we’ll cover today.
人们喜欢谈论现代生活的挑战,确实有一些,但人类历史上也从未有过比现在更好的时代。我怀疑,唯一更好的时代都在未来,这要归功于我们今天要讲的故事类型。
What a time to be alive.
多么美好的时代啊。
(1) World Model Week:
(1)世界模型周:
Runway Solaris
Runway Solaris
&
和
GWM 2
GWM 2
and
以及
World Labs Atlas
World Labs Atlas
World Model enjoyers stand up!
世界模型的爱好者们,站起来!
In March, General Intuition founder & CEO Pim DeWitte and I co-wrote an
三月,General Intuition的创始人兼CEO Pim DeWitte和我合写了一篇
essay on World Models
关于世界模型的文章
. “World models,” we wrote, “these systems that learn from watching the world and the actions taken in it — are a fundamentally new kind of foundation model. They can compute what was previously uncomputable.”
。我们写道:“世界模型,这些通过观察世界和其中的行动来学习的系统,是一种全新的基础模型。它们可以计算以前无法计算的东西。”
There are a bunch of different kinds of models that are called World Models, and in the essay we wrote that Runway’s were Video Models and World Labs’ were 3D Reconstruction and Generation Models, but both took a step closer to our definition with these releases (Runway a little more than World Labs), but whatever you call them… holy shit are these releases cool.
有许多不同类型的模型被称为世界模型,我们在文章中写道,Runway的是视频模型,World Labs的是3D重建和生成模型,但这些发布让两者都更接近我们的定义(Runway比World Labs更接近一些),但无论你怎么称呼它们……这些发布真是太酷了。
On Monday, Runway
周一,Runway
dropped Solaris
发布了Solaris
, which it calls an “Interface World Model.” Instead of an LLM writing code that a browser renders into an app, Solaris generates the app itself, frame by frame, in real time, and when you click or drag something it generates the next frame. There is no code. “The image is the application.” This is a picture is worth 1,000 words situation, so just watch the video. Software is going to get so cool.
,他们称之为“界面世界模型”。与LLM编写代码再由浏览器渲染成应用不同,Solaris实时逐帧生成应用本身,当你点击或拖动某物时,它会生成下一帧。没有代码。“图像就是应用。”这是一图胜千言的场景,所以直接看视频吧。软件会变得非常酷。
Not to be outdone, on Tuesday, Fei-Fei Li and the World Labs team
不甘示弱,周二,李飞飞和World Labs团队
showed off Atlas
展示了Atlas
, an “omni world model” trained from scratch on text, images, video, camera poses, depth maps, and 3D all at once. It generates up to a minute of 1440p video with what they call pixel-perfect camera control, meaning you hand it real camera geometry instead of typing “pan left,” it will invent a coherent hallway between two unrelated photos if you tell it they’re in the same building, and it reconstructs real scenes into 3D from anywhere between one photo and a hundred. Ehhh.. you know what? Just watch this one, too:
,一个“全能世界模型”,从头开始同时训练文本、图像、视频、相机姿态、深度图和3D。它能生成长达一分钟的1440p视频,具有他们所说的像素级完美相机控制,意味着你给它真实的相机几何而不是输入“向左平移”,如果你告诉它两张不相关的照片在同一栋建筑里,它会发明一个连贯的走廊,并且它能从一张到一百张照片之间的任何数量重建真实场景为3D。呃……你知道吗?也看看这个吧:
And
然后
then
昨天
, yesterday, Runway dropped another model -
,Runway又发布了一个模型——
GWM Wo
GWM Wo