Knowledge File / AI小生意项目库
Anthropic的Claude Opus 5成本远低于Fable 5,同时在大多数基准测试中与之持平或超越
Anthropic的Claude Opus 5在人工智能分析智能指数中得分61,领先于其他模型,但在事实准确性和高幻觉率(50%)方面存在弱点。其成本低于竞争对手Fable 5,尤其在高端性能水平上性价比突出。
SOURCE / AI小生意项目库
MIN / 9
ACCESS / 会员
POST / 2026-07-25 17:31:00
原贴
查看原文原文
According to Artificial Analysis, Anthropic's Claude Opus 5 is currently the most capable model, achieving an Intelligence Index score of 61, with particular strengths in analytical quality and knowledge-based tasks. Because Opus 5 tends to respond more frequently even when uncertain, its hallucination rate climbs to 50 percent, raising questions about reliability in high-stakes applications. At the "high" and "xhigh" performance levels, Opus 5 outperforms both Opus 4.8 and Sonnet 5 while maintaining lower costs. Anthropic's Claude Opus 5 is the most capable AI model available today, according to several benchmarks, outperforming Fable 5 while costing less. Opus 5 scored 61 on the Artificial Analysis Intelligence Index , which combines nine tests covering knowledge work, coding, scientific reasoning, and factual accuracy. That puts it just ahead of Claude Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57), and Claude Opus 4.8 (56). Artificial Analysis worked with Anthropic to test the model before its public release. In coding, Claude Opus 5 at "xhigh" paired with Claude Code shares first place on the Artificial Analysis Coding Index, which measures how well AI models handle programming tasks on their own, including finding and fixing bugs. On Terminal-Bench v2.1, which tests agents as autonomous engineers in real terminal environments, Opus 5 scored 89 percent at "max," matching the previous leader GPT-5.6 Sol. Ad For scientific reasoning, Opus 5 scored 53 percent on Humanity's Last Exam, a very difficult knowledge test covering many academic fields. That ties it with Fable 5. On CritPt, a physics benchmark from researchers at Argonne National Laboratory and UIUC, it again matches Fable 5 but trails GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra. Ad DEC_D_Incontent-1 Factual accuracy remains a weak spot: On AA-Omniscience, which tests the accuracy of a model's knowledge claims, Opus 5 improved by 7 points over Opus 4.8 but still trails Fable 5. Opus 5 also answers more often when it's uncertain, pushing its hallucination rate up 14 points to 50 percent. Epoch AI has also tested Claude Opus 5. The research institute gave it an overall Epoch Capability Index score of 159, just below Fable 5 at 161. When looking solely at the software engineering benchmarks (SWE-ECI), though, Opus 5 ties with Fable 5 at 161, outperforming GPT-5.6 Terra and Claude Opus 4.8. GPT-5.6 Sol leads in both categories. Ad Overall, this confirms that the race among frontier models is tight. No single model can pull away or claim a clear advantage. That lends weight to the argument that AI models will eventually become commoditized . The average Intelligence Index task costs $2.03 with Opus 5, less than Claude Fable 5 with fallback at $2.75. It costs more than Opus 4.8 at $1.80 and Sonnet 5 at $1.53. At the "high" and "xhigh" tiers, however, Opus 5 beats both Opus 4.8 and Sonnet 5 while costing less. Ad DEC_D_Incontent-2 Vals.ai tested Claude Opus 5 across all five reasoning tiers using Vibe Code Bench, a benchmark for programming tasks. Scores climb from 76.7 percent at "low" to 82 percent at "medium" and 89.8 percent at "high." Performance dips at the two highest tiers, with "xhigh" scoring 88.3 percent and "max" scoring 88.4 percent despite much higher costs. Ad
中文翻译
根据人工智能分析,Anthropic的Claude Opus 5是目前最强大的模型,智能指数得分为61,在分析质量和基于知识的任务方面特别强大。由于Opus 5即使在不确定时也倾向于更频繁地回应,其幻觉率攀升至50%,引发了对高可靠性应用场景的质疑。在“高”和“超高”性能水平上,Opus 5优于Opus 4.8和Sonnet 5,同时保持较低成本。Anthropic的Claude Opus 5是当前最强大的人工智能模型,根据多项基准测试,它超越Fable 5且成本更低。Opus 5在人工智能分析智能指数上得分为61,该指数综合了涵盖知识工作、编程、科学推理和事实准确性的九项测试。这使其略领先于Claude Fable 5(60)、GPT-5.6 Sol(59)、Kimi K3(57)和Claude Opus 4.8(56)。人工智能分析公司与Anthropic合作,在该模型公开发布前进行了测试。在编程方面,Claude Opus 5在“超高”模式下配合Claude Code,在人工智能分析编程指数上并列第一,该指数衡量AI模型独立处理编程任务(包括查找和修复错误)的能力。在Terminal-Bench v2.1(测试智能体在真实终端环境中作为自主工程师的能力)上,Opus 5在“最大”模式下得分为89%,与之前领先的GPT-5.6 Sol持平。在科学推理方面,Opus 5在人类最后的考试(一项涵盖多学术领域的极难知识测试)中得分为53%,与Fable 5持平。在CritPt(阿贡国家实验室和伊利诺伊大学厄巴纳-香槟分校研究人员开发的物理学基准)上,它再次与Fable 5持平,但落后于GPT-5.6 Sol、GPT-5.5 Pro和GPT-5.6 Terra。事实准确性仍是一个弱点:在AA-Omniscience(测试模型知识主张准确性的测试)上,Opus 5比Opus 4.8提高了7分,但仍落后于Fable 5。Opus 5在不确定时也更频繁地回答,将其幻觉率推高了14个百分点,达到50%。Epoch AI也测试了Claude Opus 5。该研究机构给它的总体Epoch能力指数得分为159,略低于Fable 5的161。然而,仅看软件工程基准(SWE-ECI),Opus 5以161分与Fable 5持平,优于GPT-5.6 Terra和Claude Opus 4.8。GPT-5.6 Sol在这两项中均领先。总体而言,这证实了前沿模型之间的竞争十分激烈。没有一个模型能拉开差距或宣称拥有明显优势。这支持了AI模型最终将商品化的论点。Opus 5在智能指数上的平均任务成本为2.03美元,低于带回退的Claude Fable 5的2.75美元。它高于Opus 4.8的1.80美元和Sonnet 5的1.53美元。然而,在“高”和“超高”级别上,Opus 5在成本低于Opus 4.8和Sonnet 5的同时,性能胜过它们。Vals.ai使用Vibe Code基准(一个编程任务基准)在所有五个推理级别上测试了Claude Opus 5。得分从“低”的76.7%上升到“中”的82%和“高”的89.8%。在两个最高级别上性能下降,“超高”得分为88.3%,“最大”得分为88.4%,尽管成本更高。
核心信息
Anthropic的Claude Opus 5在人工智能分析智能指数中得分61,领先于其他模型,但在事实准确性和高幻觉率(50%)方面存在弱点。其成本低于竞争对手Fable 5,尤其在高端性能水平上性价比突出。
- Anthropic的Claude Opus 5在人工智能分析智能指数中得分61,领先于其他模型,但在事实准确性和高幻觉率(50%)方面存在弱点。其成本低于竞争对手Fable 5,尤其在高端性能水平上性价比突出。
- 原贴提到:According to Artificial Analysis, Anthropic's Claude Opus 5 is currently
- 来源:the-decoder.com
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容