觉
AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-05-29 2 浏览 免费阅读

趋势解读:Anthropic ships Claude Opus 4.8 as a "modest,聚焦 Agent 工作流自动化

Anthropic发布Claude Opus 4.8,性能超越竞品,引入动态工作流和努力控制,提升Agent自动化能力。

SOURCE / AI技能杠杆 MIN / 9 ACCESS / 免费阅读 POST / 2026-05-29 05:20:09

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Anthropic has released Claude Opus 4.8, a new AI language model that the company claims outperforms competitors like OpenAI's GPT-5.5 across most benchmarks, while also communicating its own uncertainties better. Anthropic also introduces dynamic workflows that allow it to schedule tasks and launch hundreds of parallel subagents, along with a new control that lets users determine how much effort the AI should put into generating a response. API pricing remains unchanged from its predecessor, Opus 4.7, at $5 per million input tokens and $25 per million output tokens. Anthropic's latest flagship model, Claude Opus 4.8, leads most benchmarks and is designed to be more upfront about its own mistakes. Anthropic says Opus 4.8 beats both its predecessor and OpenAI's GPT-5.5 and Google's Gemini 3.1 Pro across most tested categories. On agentic coding (SWE-Bench Pro), the model hits 69.2 percent, up from 64.3 percent for Opus 4.7 and 58.6 percent for GPT-5.5. For multidisciplinary reasoning (Humanity's Last Exam), Opus 4.8 scores 49.8 percent without tools and 57.9 percent with tools, the highest marks in the field. Anthropic calls the model's improved honesty one of its most noticeable upgrades. AI models have a habit of jumping to conclusions and claiming progress that falls apart on closer look. It's a widespread problem. Ad "Early testers report that Opus 4.8 is more likely to flag uncertainties about its work and less likely to make unsupported claims," Anthropic says. The company backs that up with its own coding evaluations, where the model lets bugs slip through without comment about four times less often than Opus 4.7. Ad DEC_D_Incontent-1 The model also sets new highs on prosocial traits like supporting user autonomy. Deception attempts and other unaligned behavior are said to be at Claude Mythos levels. Details are in the Claude Opus 4.8 System Card . The first Mythos-class models are expected to roll out to all customers in the coming weeks, once all safety measures are in place, the company says. The new features Anthropic shipped alongside the model may matter more than the model update itself, which the company calls "modest but tangible." Ad The biggest is "dynamic workflows." The model can plan a task and then spin up hundreds of parallel sub-agents in a single session. Anthropic says Claude Code with Opus 4.8 can now handle codebase-wide migrations across hundreds of thousands of lines, from planning all the way to merge. The feature is available on Enterprise, Team, and Max plans. On claude.ai and in Cowork, there's now an effort control next to the model picker. It lets you decide how hard Claude works on a given response. Crank it up for deeper thinking and better results. Turn it down for faster answers that use less of your rate limit. Ad DEC_D_Incontent-2 Opus 4.8 defaults to "high." For tough tasks, Anthropic recommends "extra" (called "xhigh" in Claude Code) or "max." These modes burn more tokens, but Anthropic says higher rate limits for Claude Code users help offset that. Anthropic's advice is to just pick whatever level feels right for the task. Ad

中文翻译

Anthropic发布了Claude Opus 4.8,这是一款新的AI语言模型,公司声称其在大多数基准测试中优于OpenAI的GPT-5.5等竞争对手,同时能更好地传达自身的不确定性。Anthropic还引入了动态工作流,允许其安排任务并启动数百个并行子代理,以及一个新的控件,让用户决定AI应该在生成响应时投入多少努力。

核心信息

Anthropic发布Claude Opus 4.8,性能超越竞品,引入动态工作流和努力控制,提升Agent自动化能力。

  • Anthropic发布Claude Opus 4.8,性能全面超越前代和竞品。
  • 新引入动态工作流,可并行处理数百个子任务。
  • 强化模型坦诚度,降低错误断言和未支持声明。
  • 新增努力控制滑块,用户可调节AI响应深度与速度。

详细解读

这是什么信号?

Anthropic发布Claude Opus 4.8,虽自称为“适度但切实”的升级,但性能全面领先,并推出动态工作流和努力控制两项关键功能。这表明AI竞争从单一模型能力转向Agent工作流自动化与用户可控性的结合。

为什么重要?

动态工作流允许单会话内并行启动数百个子代理,实现代码库级迁移等复杂任务的自动化,大幅提升效率。努力控制让用户根据任务需求调节AI深度与速度,优化成本与响应时间。此外,模型坦诚度的提升减少了错误断言,增强了可靠性。

对谁有价值?

企业开发团队:可利用动态工作流重构大规模代码库;AI应用开发者:可集成努力控制以平衡性能与成本;研究者:可评估模型在复杂推理和多任务场景下的表现。

可以怎么行动?

1. 在企业计划中启用Claude Opus 4.8,测试动态工作流在代码迁移、数据处理等场景的效果。2. 根据任务复杂性调整努力控制级别,例如简单查询使用“低”模式节省成本,复杂分析使用“高”或“最高”。3. 利用模型坦诚度提升,减少事后审查工作。

风险或限制

动态工作流和“最高”努力模式会消耗更多token,增加API费用。模型虽然更坦诚,但仍可能遗漏错误,需要人工复核。目前仅有企业、团队和Max计划可用,个人用户受限。

信息差价值

信息差价值: Claude Opus 4.8的动态工作流能力是当前AI Agent领域的突破方向,多数玩家尚未实现单会话内数百子代理的并行调度。这一功能直接提升复杂任务自动化效率,信息差在于理解其规模效应和实际成本。

业务启发: 企业可立即将动态工作流应用于代码库迁移、大规模数据处理和自动化测试,缩短项目周期。努力控制则为不同部门(如客服用快速模式、研发用深度模式)提供了灵活的成本管理工具,启发企业建立基于任务价值的分级AI使用策略。

可沉淀动作: 建议团队在隔离环境中测试Opus 4.8的代码重构能力,记录动态工作流下的成功率与token消耗;同时制定公司内部的努力控制标准,并培训员工根据任务调整设置。长期可建立Agent工作流模板库,沉淀最佳实践。

参考来源

AI SUMMARY

这篇文章回答了什么

趋势解读:Anthropic ships Claude Opus 4.8 as a "modest,聚焦 Agent 工作流自动化主要讲什么?

Anthropic发布Claude Opus 4.8,性能超越竞品,引入动态工作流和努力控制,提升Agent自动化能力。

这篇文章最值得关注的要点是什么?

Anthropic发布Claude Opus 4.8,性能超越竞品,引入动态工作流和努力控制,提升Agent自动化能力。;Anthropic发布Claude Opus 4.8,性能全面超越前代和竞品。;新引入动态工作流,可并行处理数百个子任务。;强化模型坦诚度,降低错误断言和未支持声明。

这篇文章和哪些AI专题相关?

它适合放在Agent工作流、AI工具、AI超级个体专题里阅读。 关联原因:这篇内容命中「Agent、工作流」等主题信号。;这篇内容命中「自动化、模型、Claude」等主题信号。;这篇内容命中「技能」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 阶跃星辰 Step 3.7 Flash 发布,聚焦智能体效率 下一篇 用好 Coding Agent,重点是两头,尤其是开头的部分,如果一开始就走偏了后面怎么改都改不好。