AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-07-01 7 浏览 公开

Anthropic的新Claude Sonnet 5缩小了与更昂贵的Opus系列模型的差距

Anthropic发布了Claude Sonnet 5,公司称其是最具代理性的Sonnet。它能够自主制定计划并使用浏览器和终端等工具。在基准测试中,Sonnet 5全面超越了前代Sonnet 4.6,并接近更大的Opus 4.8。在真实世界的知识工作任务中,它甚至略微超过了Opus 4.8。该模型现在以介绍折扣价在所有Anthropic平台上可用,价格将在2026年8月后升至标准Sonnet费率。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-07-01 02:46:10

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Anthropic released Claude Sonnet 5, which the company calls its most agentic Sonnet yet. It can build plans on its own and use tools like browsers and terminals. In benchmarks, Sonnet 5 beats its predecessor, Sonnet 4.6, across the board and closes in on the larger Opus 4.8. On real-world knowledge work tasks, it even edges past Opus 4.8. The model is available now on all Anthropic platforms at an introductory discount, with pricing rising to standard Sonnet rates after August 2026. Anthropic released Claude Sonnet 5. In benchmarks, it closes in on the larger Opus 4.8 and even beats it in some areas. The model is available now at an introductory price. Anthropic calls it the most agentic Sonnet yet: it can build plans, grab tools like browsers and terminals, and work on its own at a level that just months ago only bigger, pricier models could pull off, according to the company. Sonnet 5 is meant to close that gap. Anthropic's published benchmarks show Sonnet 5 beating its predecessor Sonnet 4.6 in every tested category while gaining ground on the pricier Opus 4.8. On agentic coding, Sonnet 5 hits 63.2 percent on SWE-bench Pro, up from 58.1 percent for Sonnet 4.6. Opus 4.8 sits at 69.2 percent. On Terminal-Bench 2.1, Sonnet 5 pulls 80.4 percent versus Sonnet 4.6's 67.0 percent. For multidisciplinary reasoning (Humanity's Last Exam), the model reaches 57.4 percent with tools, nearly matching Opus 4.8 at 57.9 percent. On computer use (OSWorld-Verified), Sonnet 5 posts 81.2 percent compared to 78.5 percent for its predecessor. Ad On the knowledge work benchmark GDPval-AA v2, which tests AI on real-world knowledge tasks , Sonnet 5 actually beats the larger Opus 4.8, scoring 1,618 to Opus's 1,615. Anthropic says feedback from early-access partners told the same story. Sonnet 5 acts far more agentically than previous versions, showing up in things like how it handles search tasks. Ad DEC_D_Incontent-1 Lately, Anthropic has been making news for models it can't ship. The US government is blocking the company's two most capable models, Mythos 5 and Fable 5 , over cybersecurity concerns. That context hangs over the Sonnet 5 launch. Anthropic is clearly eager to get ahead of any similar worries. The model wasn't trained on cybersecurity tasks, the company says, and in tests for risky capabilities like writing software exploits, it scores far below both Opus 4.8 and Mythos 5. Sonnet 5 does score a bit higher than its predecessor on these tasks, though. So Anthropic has switched on cyber safeguards by default. They flag and block risky cyber usage in real time, on par with the protections already in place for Claude Opus 4.7 and 4.8. They're dialed back compared to Fable 5's guardrails, which users complained about almost immediately . Anthropic says it views the overall cybersecurity risk from Sonnet 5 as low. Ad On the safety front, the model does a better job turning down malicious requests and fending off prompt injection attacks than Sonnet 4.6, according to Anthropic. Hallucinations and sycophantic behavior , the tendency to just agree with whatever the user says, are down as well. Anthropic's full safety evaluation is in the Claude Sonnet 5 System Card . Claude Sonnet 5 is live now on all plans. It's the new default for Free and Pro users, and Max, Team, and Enterprise subscribers can access it too. Developers can plug it into Claude Code and the Claude Platform. On the API side, it goes by "claude-sonnet-5" . The training cutoff is January 2026, with a one-million-token context window. Ad DEC_D_Incontent-2 Until August 31, 2026, Anthropic is charging $2 per million input tokens and $10 per million output tokens. After that , prices jump to $3 and $15, which is what previous Sonnet models cost. Ad

中文翻译

Anthropic发布了Claude Sonnet 5,公司称其是最具代理性的Sonnet。它能够自主制定计划并使用浏览器和终端等工具。在基准测试中,Sonnet 5全面超越了前代Sonnet 4.6,并接近更大的Opus 4.8。在真实世界的知识工作任务中,它甚至略微超过了Opus 4.8。该模型现在以介绍折扣价在所有Anthropic平台上可用,价格将在2026年8月后升至标准Sonnet费率。Anthropic发布了Claude Sonnet 5。在基准测试中,它接近更大的Opus 4.8,并在某些领域甚至超越了它。该模型现在以介绍价提供。Anthropic称其是最具代理性的Sonnet:据公司称,它可以制定计划、获取浏览器和终端等工具,并自主工作,达到几个月前只有更大、更昂贵的模型才能做到的水平。Sonnet 5旨在缩小这一差距。Anthropic发布的基准测试显示,Sonnet 5在每个测试类别中都超过了前代Sonnet 4.6,同时缩小了与更昂贵的Opus 4.8的差距。在代理编码方面,Sonnet 5在SWE-bench Pro上达到63.2%,而Sonnet 4.6为58.1%,Opus 4.8为69.2%。在Terminal-Bench 2.1上,Sonnet 5达到80.4%,而Sonnet 4.6为67.0%。在多学科推理(Humanity's Last Exam)中,该模型使用工具达到57.4%,几乎与Opus 4.8的57.9%持平。在计算机使用(OSWorld-Verified)方面,Sonnet 5达到81.2%,而前代为78.5%。在知识工作基准测试GDPval-AA v2上,该测试评估AI在真实世界知识任务上的表现,Sonnet 5实际上超过了更大的Opus 4.8,得分1,618对Opus的1,615。Anthropic表示,早期访问合作伙伴的反馈也证实了这一点。Sonnet 5的代理性远高于以前的版本,体现在它处理搜索任务等方面。Ad DEC_D_Incontent-1 最近,Anthropic因无法发布的模型而登上新闻。美国政府出于网络安全考虑,阻止了该公司两个最强大的模型Mythos 5和Fable 5。这一背景给Sonnet 5的发布蒙上了阴影。Anthropic显然希望避免类似的担忧。该公司表示,该模型未在网络安全任务上进行训练,并且在编写软件漏洞等危险能力的测试中,得分远低于Opus 4.8和Mythos 5。不过,Sonnet 5在这些任务上的得分比前代略高。因此,Anthropic默认开启了网络保护措施。它们实时标记和阻止危险的网络使用,与Claude Opus 4.7和4.8已有的保护水平相当。相比于Fable 5的护栏(用户几乎立即抱怨),这些保护有所减弱。Anthropic表示,Sonnet 5的整体网络安全风险较低。Ad 在安全方面,据Anthropic称,该模型在拒绝恶意请求和防御提示注入攻击方面比Sonnet 4.6做得更好。幻觉和谄媚行为(即倾向于同意用户所说的任何话)也有所减少。Anthropic的完整安全评估见Claude Sonnet 5系统卡。Claude Sonnet 5现已对所有计划上线。它是Free和Pro用户的新默认模型,Max、Team和Enterprise订阅者也可以使用。开发者可以将其接入Claude Code和Claude平台。在API方面,它被称为“claude-sonnet-5”。训练截止日期为2026年1月,上下文窗口为一百万个token。Ad DEC_D_Incontent-2 截至2026年8月31日,Anthropic收取每百万输入token 2美元,每百万输出token 10美元。之后,价格将升至3美元和15美元,这是之前Sonnet模型的价格。Ad 来源:The Decoder 分数:82

核心信息

Anthropic发布了Claude Sonnet 5,公司称其是最具代理性的Sonnet。它能够自主制定计划并使用浏览器和终端等工具。在基准测试中,Sonnet 5全面超越了前代Sonnet 4.6,并接近更大的Opus 4.8。在真实世界的知识工作任务中,它甚至略微超过了Opus 4.8。该模型现在以介绍折扣价在所有Anthropic平台上可用,价格将在2026年8月后升至标准Sonnet费率。

  • Anthropic发布了Claude Sonnet 5,公司称其是最具代理性的Sonnet。它能够自主制定计划并使用浏览器和终端等工具。在基准测试中,Sonnet 5全面超越了前代Sonnet 4.6,并接近更大的Opus 4.8。在真实世界的知识工作任务中,它甚至略微超过了Opus 4.8。该模型现在以介绍折扣价在所有Anthropic平台上可用,价格将在2026年8月后升至标准Sonnet费率。
  • 原贴提到:Anthropic released Claude Sonnet 5, which the company calls its most age
  • 来源:the-decoder.com

详细解读

这是什么信号?

Anthropic发布Claude Sonnet 5是一个重要信号:AI模型正在快速向“代理化”演进,同时性能差距在缩小。Sonnet 5作为中等价位的模型,在多项基准上逼近甚至超越高价Opus系列,表明AI能力正从“大而贵”向“小而精”迁移。此外,美国政府因网络安全问题阻止Mythos 5和Fable 5发布,凸显了AI安全监管的加强,而Sonnet 5特意弱化网络安全能力并默认开启防护,表明Anthropic在努力平衡能力与合规。

为什么重要?

对于AI行业,Sonnet 5的发布意味着企业和开发者可以以更低成本获得接近顶级的代理能力。SWE-bench Pro从58.1%提升到63.2%,Terminal-Bench从67.0%提升到80.4%,这些进步直接转化为实际效率提升。同时,安全问题的处理方式(降低风险能力、加强护栏)可能成为行业标准。

对谁有价值?

• 开发者和AI产品团队:可以立即用Sonnet 5构建更自主的代理应用,如自动化编码、浏览器操作等,且成本可控。
• 企业决策者:需要评估模型选择,Sonnet 5在知识工作(GDPval-AA v2)上甚至略胜Opus 4.8,适合替代昂贵模型。
• AI安全研究者:Sonnet 5的安全评估和护栏设计提供了新的案例。

可以怎么行动?

• 试用API:在2026年8月前享受折扣,集成到现有工作流中测试代理性能。
• 调整AI策略:将非关键任务从Opus迁移到Sonnet 5以节省成本;对于需要代理能力但预算有限的场景,首选Sonnet 5。
• 关注合规:部署时需确认Sonnet 5的默认安全设置是否满足企业要求,特别是涉及敏感数据时。

风险或限制

• 网络安全风险:尽管声称风险低,但Sonnet 5在危险任务上得分高于前代,且默认防护可能不够严格,需要人为监控。
• 性能上限:在SWE-bench Pro等任务上仍落后Opus 4.8约6个百分点,顶尖任务仍需高端模型。
• 监管不确定性:美国政府可能后续将Sonnet 5纳入审查范围,影响长期可用性。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《Anthropic的新Claude Sonnet 5缩小了与更昂贵的Opus系列模型的差距》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

上一篇 AIHOT 日报参考 2026-07-01 下一篇 GitHub如何维护开源依赖的合规性