【文章标题】:GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour
【文章标题】:GPT-6 Astra:每小时成本低于6美元的可雇佣自动化AI工程师

【文章正文】:
GPT-6 Astra
GPT-6 Astra

, the first
,作为首个

Stargate
”星际之门”

and
且

lightly looped
轻量级循环

supermodel from OpenAI,
的OpenAI超级模型,

launched today
于今日发布

, cleanly beating
,在多维度碾压

Fable 5.1
Fable 5.1

on many metrics including completely saturating the hardest versions of
其指标包括完全攻克最艰难版本的

FrontierMath
FrontierMath数学测试(97.6%)

(97.6%) and
与

ARC-AGI-3
ARC-AGI-3测试(99.9%)。

(99.9%). Lots of demos will focus on typical talk tracks like the
大量演示将聚焦常规场景:从

computer use
计算机操作

to the
到

Pokemon playing
《宝可梦》游戏

to
再到

Blender
Blender建模

to the
乃至

scientific
科学

and
与

cybersafety
网络安全

benchmarks (
基准测试(

system card
系统卡片

). Greg says
)。Greg宣称

AGI is here
”AGI已降临”

, and Jakub says it is
Jakub则认为这是

finally the Automated AI Research Intern
”终于等来的自动化AI研究实习生”

he wanted.
他梦寐以求的。

We aren’t qualified to talk about those, but we got early access and threw it at every practical, real-life task we could think of. After burning
我们虽无资格评价这些,但通过早期实测将其投入所有能想到的现实任务后——在消耗

over 20B tokens of Astra
超200亿token后

, we can confirm the most surprising finding:
可确认最惊人的发现:

GPT-6 Astra
GPT-6 Astra

is one of a new class of models
属于新一代模型

1
注1

that are fully capable AI Engineers in their own right
其本身即具备完整AI工程师能力

. They now help you
现可协助您

choose and train models
选择训练模型

,
、

label data
标注数据

(both helping you label and then using your labels for active learning, like
(既辅助标注又利用标注进行主动学习,如

SAM
SAM模型

),
)、

keep pipelines saturated
维持流水线饱和

,
、

instrument and read logs
埋点与日志解读

,
、

deploy and debug entire systems
一键部署调试完整系统

in one shot, fan out and
分派并

command and eval subagents
指挥评估子代理

(including agents running other models), and keep coherence over
(含运行其他模型的代理),且在单线程

billions
数十亿

of tokens of a single agent thread.
token量级保持一致性。

Raising Your Ambitions
雄心升级

We’ve written before about
我们曾撰文探讨

the high-return activity of raising your aspirations for LLMs
”提升对LLM期望值的高回报活动”

.
。

Our experience has made us exponentially more ambitious than we have ever been
实测使我们的野心呈指数级增长

. Over the past month, we went from prompting humans for a fun “
过去一个月,我们从举办人类参与的趣味”

Kill My SaaS
杀死我的SaaS”竞赛

2
注2

” competition
,到构建

, to building a
十余款内部/个人工具

dozen internal/personal tools
,包括

, including
4款原需付费的SaaS工具

4 previously paid SaaS tools
,彻底

, fully
重构个人网站

redesigned my personal site
,开发不完善但可用的

, made an incomplete but functional
GitHub+Vercel替代方案

replacement of GitHub + Vercel
,为

, trained game AI for
策略棋盘游戏训练AI(合法走法超围棋万倍)

a strategy board game with 10,000x more legal moves than Go
,在个人财务整理中节省数万美元

, saved tens of thousands of dollars in personal finance cleanups,
,重新出版旧书

republished my old book
并同步有声书与印刷版

with synced audiobook audio and printed physical editions, and even
甚至即将启动

more
更

ambitious
宏大的

projects we will launch soon.
项目。

The $6 an hour number might sound surprising, but that’s exactly what we saw in
时薪6美元或令人惊讶,但这正是我们

our testing
测试所见

  • 33 tokens per second at a max $50 per million token rate. Given that Astra is more token efficient than Sol and Fable (
    ——每秒33token,每百万token最高50美元。鉴于Astra比Sol和Fable更具token效率(

independently confirmed by Artificial Analysis
Artificial Analysis独立证实

), it often means that Astra is simultaneously also the best fast-and-smart model you can buy (assuming our preview latency holds for GA), outside of
),常意味着Astra同时是能买到的最佳快智模型(假设预览版延迟在正式版保持),仅次于

Spark 1.3
Spark 1.3

.
。

see logs
参见日志

Managing fleets of subagents (individually tweaked, bounded concurrency)
子代理舰队管理(独立调参,有限并发)

Now of course, if you just throw on Astra at Ultra you’re gonna burn through a lot more than $6 per hour…. because it is so dang good at parallelizing. Depending on the task in practice we were often ramping up
当然,若直接启用Ultra版Astra,时耗将远超6美元…因其并行能力极强。实测中我们常启动

between 20-50 agents in parallel
20-50个并行代理

, of course all managed by one main Astra agent.
(均由单个Astra主代理管理)。

Monitoring its own runs, starting and stopping waves
自主监控运行、启停任务波次

This is basically what you would pay a junior AI Engineer to do — babysitting runs, staring at data, finding issues, fixing, rerunning, ad infinitum. You could hire someone at 1000 a day, or you can hire GPT-6 for $100 over 2 days to do this.
这本质是初级AI工程师的工作——看守运行、盯数据、查问题、修复、重试,循环往复。您可日薪200-1000美元雇人,或两天100美元雇佣GPT-6完成。

Making model benchmarks, handling budgets, making estimates, scaling up runs, getting human ratings
制作模型基准、处理预算、预估成本、扩展运行、获取人工评分

Because of course you need all these capabilities to run your own AI engineering program, because of course OpenAI already uses GPT-6 to do this internally…
因运营AI工程项目必然需要这些能力——毕竟OpenAI内部早已用GPT-6做这些…

example
示例

here
此处

Or you can get Astra to trivially whip up your own
或让Astra轻松搭建

personal Arena.ai clone
个人版Arena.ai

for tuning your prompts, picking models for your task, or aligning yrou own preference model!
用于提示词调优、任务模型选择或对齐偏好模型!

The overall conclusion you should have is that
核心结论应是:

OpenAI have clearly trained a model that is capable of automating much of their own AI Engineering
OpenAI显然训练出了能自动化其大部分AI工程的模型

, and it is finally time that you learn to exploit Astra- and Fable-class models and be far,
现在正是学习利用Astra与Fable级模型之时,且应对智能体能力抱持

far more unreasonable
远超常理的

with your own expectations of what you can do with agents now.
期待。

1
注1

We are
我们正对

running similar work
Grok、Fable等前沿模型

on Grok, Fable and other similar frontier models but OpenAI was most generous with trial limits so this gets the writeup - but the agentic coding patterns discussed here will likely apply to all such late 2026 frontier models.
开展类似测试,但OpenAI试用限额最慷慨故成文——本文讨论的代理编码模式可能适用于所有2026年末代前沿模型。

2
注2

Many of you are waiting to hear results… sorry for the radio silence! we got… busy! We will announce winners and reimbursements and best attempts.
许多读者等待结果…抱歉静默!我们…太忙了!将公布获胜者、补偿与最佳尝试。

🔗 知识库双向关联