AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-07-25 3 浏览 公开

Claude Opus 5 发布

Anthropic 发布新模型 Claude Opus 5,价格仅为 Claude Fable 5 的一半,性能领先,具有主动性和安全考虑。用户尚未测试,但整体评价积极。

SOURCE / 全球热点解读 MIN / 4 ACCESS / 公开 POST / 2026-07-25 07:48:50

原贴

查看原文
作者:Simon Willison 来源站点:simonwillison.net 原贴时间:

原文

Introducing Claude Opus 5 I've been offline kayaking with sea otters for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its paces yet. The buzz is positive, and Anthropic's description of it as a "thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price" sounds promising. It's currently leading the Artificial Analysis leaderboard , in front of even Fable 5. It's priced the same as Opus 4.8, and continues to offer a "fast mode" at twice the cost of the base model. Based on this anecdote in the release post it sounds like it might be relentlessly proactive : On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly viewthe drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It's better at finding vulnerabilities but has deliberately not been trained on how to exploit them. Hopefully this means the US government won't shut it down! As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats. Anthropic have published a prompting guide for Claude Opus 5 . The first pelican I got was missing the bicycle wheels; the second attempt was better. Tags: ai , generative-ai , llms , anthropic , claude , llm-release

中文翻译

我今天大部分时间都在离线与海獭划皮划艇,所以还没能测试 Anthropic 的新模型 Claude Opus 5。但热议是正面的,Anthropic 将其描述为“一款深思熟虑且主动的模型,以一半的价格接近 Claude Fable 5 的前沿智能”,这听起来很有希望。它目前在 Artificial Analysis 排行榜上领先,甚至超过了 Fable 5。它的定价与 Opus 4.8 相同,并继续提供“快速模式”,价格为基础模型的两倍。

根据发布文章中的一段轶事,它可能非常主动:在一项 Frontier-Bench 任务中,Opus 5 被给予一张机器零件的图纸,并被要求编写代码将其重建为 3D FreeCAD 模型。然而,模型被故意没有提供直接查看图纸的方式。Opus 5 的回应是编写自己的计算机视觉管道,从原始像素中提取几何形状,然后重建了整个机器零件。

它在发现漏洞方面有所改进,但刻意没有被训练如何利用漏洞。希望这意味着美国政府不会关闭它!与其前身 Opus 4.8 一样,我们刻意避免对 Opus 5 进行网络安全任务训练。然而,由于变得更通用,该模型在这些任务上仍有显著提升,接近 Mythos 5 发现网络安全漏洞的能力。但在利用这些漏洞方面——即将漏洞转化为实质性网络威胁——它仍远远落后于 Mythos 5。

核心信息

Anthropic 发布新模型 Claude Opus 5,价格仅为 Claude Fable 5 的一半,性能领先,具有主动性和安全考虑。用户尚未测试,但整体评价积极。

  • Anthropic 发布新模型 Claude Opus 5,价格仅为 Claude Fable 5 的一半,性能领先,具有主动性和安全考虑。用户尚未测试,但整体评价积极。
  • 原贴提到:Introducing Claude Opus 5 I've been offline kayaking with sea otters for
  • 来源:simonwillison.net

详细解读

这是什么信号?

Anthropic 发布 Claude Opus 5,定位为性能接近旗舰模型 Fable 5 但价格减半的“高性价比前沿模型”。它在基准测试中领先,并展现出惊人的主动性(自主构建视觉管线解决任务)。同时,Anthropic 刻意限制其网络攻击利用能力,表明安全考量仍是关键设计原则。

为什么重要?

1. 价格-性能拐点:Opus 5 以 Opus 4.8 的价格提供接近 Fable 5 的能力,可能加速高端 AI 的普及。
2. 主动性范本:模型自主编写计算机视觉管线的行为,标志着 AI 从“被动响应”向“主动规划”演进。
3. 安全平衡术:在能力增强的同时刻意抑制有害用途,为行业树立负责任的发布标杆。

对谁有价值?

企业开发者:可用更低成本获得接近顶尖的模型能力,尤其适合复杂推理和自动化任务。
AI 安全研究者:Anthropic 的安全策略(训练时不包含利用漏洞)提供了可研究的案例。
AI 应用创业者:Opus 5 的主动特性可催化新应用场景(如自主代码修复、逆向工程)。

可以怎么行动?

1. 立即试用 Opus 5 API,对比与 Fable 5 在自身业务场景中的表现。
2. 探索主动行为场景:设计需要多步推理和工具使用的任务,测试模型自主性。
3. 关注安全边界:基于 Anthropic 的提示指南(见原文链接)构建应用,避免触发限制。

风险或限制

1. “主动”可能带来不可预测性:模型自主行为偏离用户意图时难以控制。
2. 安全克制意味着在网络安全等领域的性能上限,不适合需要漏洞利用能力的场景。
3. 基准测试与实际业务存在差距,需独立验证效果。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 simonwillison.net 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《Claude Opus 5 发布》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

上一篇 AIHOT 日报参考 2026-07-25 下一篇 英伟达、微软和Meta联合警告:应避免对开放权重模型过度监管