【文章标题】:GPT‑6 Astra

【文章正文】: GPT‑6 Astra

GPT-6 Astra is “rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS” - I’ve not tried it yet myself, so I don’t have a great deal to say about it yet. GPT-6 Astra”今日开始向部分机构限时推出,未来几天将逐步开放给所有ChatGPT Plus、Pro、Business和Enterprise用户,同时通过OpenAI API和AWS平台提供”——我本人尚未体验,因此暂时无法详述。

It’s going to be API priced at the same rate as Claude Fable 5 and 5.1: 50/million output. This is clearly OpenAI’s Fable competitor, and appears to score higher than Fable on most of OpenAI’s self-reported benchmarks. 其API定价与Claude Fable 5和5.1保持一致:每百万次输入10美元,每百万次输出50美元。这显然是OpenAI对标Fable的产品,在OpenAI自测的大部分基准测试中得分更高。

Most impressively, Astra scores 99.9% on the recent (released in March) 最令人印象深刻的是,Astra在三月份最新发布的

ARC-AGI 3 benchmark ARC-AGI 3基准测试中取得99.9%的分数

  • though notably Fable 5 does not yet have a published result, and the ——尽管值得注意的是Fable 5尚未公布测试结果,且

ARC-AGI blog notes ARC-AGI博客指出

that the 99.9% score was achieved for 26K. 99.9%的分数是通过OpenAI定制版”Provider Adapter harness”以1.9万美元成本达成,而标准ARC-AGI测试框架花费2.6万美元仅获得62.7%分数。

The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. Provider Adapter框架能在请求间保留不透明推理状态,并对长对话进行压缩存储,使模型能复用先前工作成果。

Unsurprisingly, given 鉴于

the recent Hugging Face incident 近期Hugging Face安全事件

, Astra is a beast at security tasks. It scores 100% on ExploitBench (GPT-5.6 Sol got 78.5%), 42.4% on ExploitGym (Sol got 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering compared to Sol’s 68.7%. Astra在安全任务表现堪称怪兽级:ExploitBench测试100%满分(GPT-5.6 Sol仅78.5%),ExploitGym测试42.4%(Sol为30.3%),SRE-Bench二进制逆向工程测试四次尝试内达成99.2%(Sol仅68.7%)。

It’s also better at long context: on OpenAI’s eight-needle benchmark it got 100% at 256K–512K tokens and 96.3% at 512K–1M tokens. OpenAI may have vanquished one of the ongoing challenges with long context processing. 长上下文处理同样出色:在OpenAI八针测试中,256K–512K标记长度获得100%准确率,512K–1M标记长度达96.3%。OpenAI可能已攻克长上下文处理的持续挑战。

It doesn’t win at everything though. 但并非全面领先。

Artificial Analysis Artificial Analysis指出

note that Astra is still beaten by Fable on their Intelligence Index: Astra在其智力指数上仍落后于Fable:

Sits beside GPT-5.6 Sol in Intelligence 智力指数与GPT-5.6 Sol持平

: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max). GPT-6 Astra指数得分61分与GPT-5.6 Sol持平,比Claude Fable 5.1(带后备方案的最高分)低5分,也落后于Meta新发布的Muse Spark 1.3(最高分)。

It did better on their Coding Agent Index: 在编程代理指数表现更佳:

Leads Coding Agent Index cost efficiency frontier 占据编程代理指数成本效率前沿

: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score. 全力运行时,GPT-6 Astra成本与GPT-5.6 Sol(最高配置)相当,但指数得分高出2分。单任务成本不到Claude Fable 5的一半,却能获得相同分数。

I’ll write more about Astra once I get access to it. The API model label once it rolls out will be 待获得访问权限后我将详述Astra。其API模型标签确定为

gpt-6-astra gpt-6-astra

Via 消息来源

Hacker News Hacker News

Tags: 标签:

ai 人工智能

,

openai OpenAI

,

generative-ai 生成式AI

,

llms 大语言模型

,

llm-release 大模型发布

🔗 知识库双向关联