AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-26 0 浏览 会员

Anthropic的Opus 5在衡量真实智能的基准测试中大幅超越Fable 5和GPT-5.6 Sol

Anthropic的Claude Opus 5在ARC-AGI-3基准测试中得分30.2%,是OpenAI GPT-5.6 Sol(Max)此前创下的7.8%纪录的近四倍。ARC Prize团队认为,这一领先源于更强的逻辑推理能力,使其能够在陌生环境中进行更自主的探索和规划。测试中,Opus 5表现出前所未有的行为,包括将任务转化为代数符号并独立构建反射方程,同时解决了五个此前未被解决的测试环境。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-26 17:43:02

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The ARC Prize team attributes the lead to genuinely stronger logical reasoning that enables more autonomous exploration and planning in unfamiliar environments. During testing, Opus 5 displayed behavior not previously seen from an AI model, including translating tasks into algebraic notation and independently formulating reflection equations, while also solving five previously unsolved environments. The creators of the ARC-AGI benchmark say Anthropic's Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning. The model scored 30.2 percent on ARC-AGI-3, making it the new leader. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level. That also puts it ahead of Anthropic's "Fable-class" models, which hit around 20 percent according to ARC Prize. ARC Prize's analysis credits the lead to stronger logical reasoning, "which enables more autonomous exploration, planning, and execution across unfamiliar environments." During testing, Opus 5 also showed behavior that researchers hadn't seen from a model before. It translated tasks into algebraic notation and independently formulated reflection equations for the first time. Ad Six of the 25 public demo environments have now been solved. The full results , replays , and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark , Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Both results match previous top scores, though at slightly higher costs, according to ARC Prize. Ad DEC_D_Incontent-1 ARC-AGI-3 measures how well AI models solve new tasks they didn't encounter during training, including ones humans can usually handle with ease. The current version works like a game. The model must infer the rules of an interactive environment, plan its actions, and carry them out step by step. This tests general reasoning rather than stored knowledge. Some AI systems may have already passed the benchmark, but they rely on extra software known as a harness . Official scores count only the language model's own performance. ARC Prize argues that future AGI systems shouldn't need outside help to solve new tasks. Opus 5 would likely score even higher if used within Claude Code. Ad Anthropic hasn't explained the gain, but targeted data labeling and reinforcement learning are plausible factors. Unlike earlier models, Opus 5 was developed after ARC-AGI-3 and its format became public. That may have let Anthropic target the benchmark's skills and puzzle formats, though it doesn't show the company trained on the exact tasks. Annotators could have labeled reasoning traces, useful actions, failed attempts, and recovery steps from similar puzzles. Reinforcement learning could then reward exploration, planning, rule discovery, and self-correction. Tests on Witness , Guanghan Ning's private benchmark for interactive puzzle games, point to narrower gains. Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5 while improving far less over Opus 4.8 than it did on ARC-AGI-3. On a puzzle built around common mechanics, Opus 5 stated the hidden rules before making its first move. But when a game combined rules in a less familiar way, it scored below Opus 4.8. Ning says that pattern fits training on genre-specific data, though Witness can't identify what data Anthropic used. Ad DEC_D_Incontent-2 Greg Kamradt, one of the researchers behind ARC-AGI-3 , said the results don't rule out broader reasoning gains. The conventional puzzle may not test novelty because it uses familiar mechanics, while weaker performance on one unusual game could be an isolated regression. Kamradt noted that Opus 4.8 also outperformed Opus 5 on some ARC-AGI-3 games despite trailing it by a wide margin overall

中文翻译

Anthropic的Claude Opus 5在ARC-AGI-3基准测试中得分30.2%,是OpenAI GPT-5.6 Sol(Max)此前创下的7.8%纪录的近四倍。ARC Prize团队将这一领先归因于真正更强的逻辑推理能力,使其能够在陌生环境中进行更自主的探索和规划。测试中,Opus 5表现出此前AI模型从未出现过的行为,包括将任务转化为代数符号并独立构建反射方程,同时解决了五个此前未被解决的测试环境。

ARC-AGI基准的创建者表示,Anthropic的Claude Opus 5在ARC-AGI-3上的巨大领先源于真正更好的推理能力。该模型在ARC-AGI-3上得分30.2%,成为新领先者。此前的纪录是OpenAI GPT-5.6 Sol(Max)创下的7.8%。Opus 5解决了五个此前未解决的环境,其中四个达到或超过人类水平。这也使其领先于Anthropic的“Fable-class”模型,后者据ARC Prize得分约为20%。ARC Prize的分析将领先归因于更强的逻辑推理能力,“这使其能够在陌生环境中进行更自主的探索、规划和执行。”测试中,Opus 5还表现出研究人员此前从未见过的行为。它首次将任务转化为代数符号,并独立构建反射方程。

25个公开演示环境中的六个现已得到解决。完整结果、回放和基准测试代码已公开。在较旧的ARC-AGI-2基准上,Opus 5得分为90.4%,在ARC-AGI-1上达到97.5%。据ARC Prize,两项结果均与之前最高分持平,但成本略高。

ARC-AGI-3衡量AI模型解决训练中未遇到的新任务的能力,包括人类通常能轻松处理的任务。当前版本如同游戏。模型必须推断交互环境的规则,规划行动,并逐步执行。这测试的是通用推理而非存储的知识。一些AI系统可能已经通过了基准,但它们依赖称为“harness”的外部软件。官方得分仅计入语言模型自身的表现。ARC Prize认为,未来的AGI系统不应需要外部帮助来解决新任务。如果在Claude Code中使用,Opus 5的得分可能会更高。

Anthropic尚未解释这一进步,但定向数据标注和强化学习是可能的因素。与早期模型不同,Opus 5是在ARC-AGI-3及其格式公开后开发的。这可能让Anthropic定位基准的技能和谜题格式,尽管并未显示该公司在确切任务上训练。标注者可能标记了类似谜题的推理轨迹、有用行为、失败尝试和恢复步骤。强化学习随后可能奖励探索、规划、规则发现和自纠错。

在Witness(广汉宁的交互式谜题游戏私有基准)上的测试指向更窄的进步。Opus 5得分为43.4,与Kimi K3和Fable 5统计上持平,相比Opus 4.8的提升远小于其在ARC-AGI-3上的提升。在一个围绕常见机制构建的谜题上,Opus 5在首次行动前就陈述了隐藏规则。但当游戏以不太熟悉的方式组合规则时,其得分低于Opus 4.8。宁表示,这种模式符合在特定类型数据上的训练,尽管Witness无法识别Anthropic使用了哪些数据。

ARC-AGI-3的研究者之一Greg Kamradt表示,结果并不排除更广泛的推理进步。传统谜题可能无法测试新颖性,因为它使用熟悉的机制,而在一个不寻常游戏上的较弱表现可能是孤立的回归。Kamradt指出,尽管Opus 4.8在总体上大幅落后于Opus 5,但在某些ARC-AGI-3游戏上其表现仍优于Opus 5。

核心信息

Anthropic的Claude Opus 5在ARC-AGI-3基准测试中得分30.2%,是OpenAI GPT-5.6 Sol(Max)此前创下的7.8%纪录的近四倍。ARC Prize团队认为,这一领先源于更强的逻辑推理能力,使其能够在陌生环境中进行更自主的探索和规划。测试中,Opus 5表现出前所未有的行为,包括将任务转化为代数符号并独立构建反射方程,同时解决了五个此前未被解决的测试环境。

  • Anthropic的Claude Opus 5在ARC-AGI-3基准测试中得分30.2%,是OpenAI GPT-5.6 Sol(Max)此前创下的7.8%纪录的近四倍。ARC Prize团队认为,这一领先源于更强的逻辑推理能力,使其能够在陌生环境中进行更自主的探索和规划。测试中,Opus 5表现出前所未有的行为,包括将任务转化为代数符号并独立构建反射方程,同时解决了五个此前未被解决的测试环境。
  • 原贴提到:Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 Cursor 的智能体集群表明,当前沿模型规划工作时,更便宜的模型可以处理大部分编码 下一篇 马斯克:可直接与 Grok Build 对话