【帖子标题】:Terminal Bench 4.0 just dropped, GLM-5.3 is at the same level as Fable 5, accounting for margin of error 【帖子标题】:Terminal Bench 4.0 刚刚发布,GLM-5.3 与 Fable 5 处于同一水平,误差范围内

【帖子正文】: Announcement: 【帖子正文】: 公告:

https://www.tbench.ai/news/terminal-bench-4.0 https://www.tbench.ai/news/terminal-bench-4.0

Leaderboard: 排行榜:

https://www.tbench.ai/ https://www.tbench.ai/

Imo the best aspect in their announcement is their focus on rapidly iterating on TerminalBench to keep the pace up with new model releases to fight benchmark saturation. 我认为他们公告中最棒的一点是专注于快速迭代 TerminalBench,以跟上新模型发布的速度,对抗基准测试饱和。

On a similar note, what cheaper/smaller alternatives are there to benchmarking coding agents or your own harness? Large benchmarks like this take 5-10B tokens, which is not economically/computationally feasible for the vast majority of us. 类似地,对于基准测试编码代理或你自己的测试工具,有哪些更便宜/更小的替代方案?像这样的大型基准测试需要 50-100 亿 tokens,这对我们绝大多数人来说在经济/计算上都是不可行的。

I’d love to objectively measure how my skills/harness/tools/etc change token usage and success probability on general coding tasks, there has to be a way to do this to at least give an idea or general direction, without requiring billions of tokens for each run. 我很想客观地衡量我的技能/测试工具/工具等如何改变通用编码任务中的 token 使用和成功概率,一定有办法做到这一点,至少能给出一个想法或大致方向,而不需要每次运行都消耗数十亿 tokens。