Knowledge File / AI技能杠杆
趋势解读:Anthropic ships Claude Opus 4.8 as a "modest,聚焦 Agent 工作流自动化
Anthropic发布Claude Opus 4.8,性能超越竞品,引入动态工作流和努力控制,提升Agent自动化能力。
SOURCE / AI技能杠杆
MIN / 9
ACCESS / 会员
POST / 2026-05-29 05:20:09
原贴
查看原文原文
Anthropic has released Claude Opus 4.8, a new AI language model that the company claims outperforms competitors like OpenAI's GPT-5.5 across most benchmarks, while also communicating its own uncertainties better. Anthropic also introduces dynamic workflows that allow it to schedule tasks and launch hundreds of parallel subagents, along with a new control that lets users determine how much effort the AI should put into generating a response. API pricing remains unchanged from its predecessor, Opus 4.7, at $5 per million input tokens and $25 per million output tokens. Anthropic's latest flagship model, Claude Opus 4.8, leads most benchmarks and is designed to be more upfront about its own mistakes. Anthropic says Opus 4.8 beats both its predecessor and OpenAI's GPT-5.5 and Google's Gemini 3.1 Pro across most tested categories. On agentic coding (SWE-Bench Pro), the model hits 69.2 percent, up from 64.3 percent for Opus 4.7 and 58.6 percent for GPT-5.5. For multidisciplinary reasoning (Humanity's Last Exam), Opus 4.8 scores 49.8 percent without tools and 57.9 percent with tools, the highest marks in the field. Anthropic calls the model's improved honesty one of its most noticeable upgrades. AI models have a habit of jumping to conclusions and claiming progress that falls apart on closer look. It's a widespread problem. Ad "Early testers report that Opus 4.8 is more likely to flag uncertainties about its work and less likely to make unsupported claims," Anthropic says. The company backs that up with its own coding evaluations, where the model lets bugs slip through without comment about four times less often than Opus 4.7. Ad DEC_D_Incontent-1 The model also sets new highs on prosocial traits like supporting user autonomy. Deception attempts and other unaligned behavior are said to be at Claude Mythos levels. Details are in the Claude Opus 4.8 System Card . The first Mythos-class models are expected to roll out to all customers in the coming weeks, once all safety measures are in place, the company says. The new features Anthropic shipped alongside the model may matter more than the model update itself, which the company calls "modest but tangible." Ad The biggest is "dynamic workflows." The model can plan a task and then spin up hundreds of parallel sub-agents in a single session. Anthropic says Claude Code with Opus 4.8 can now handle codebase-wide migrations across hundreds of thousands of lines, from planning all the way to merge. The feature is available on Enterprise, Team, and Max plans. On claude.ai and in Cowork, there's now an effort control next to the model picker. It lets you decide how hard Claude works on a given response. Crank it up for deeper thinking and better results. Turn it down for faster answers that use less of your rate limit. Ad DEC_D_Incontent-2 Opus 4.8 defaults to "high." For tough tasks, Anthropic recommends "extra" (called "xhigh" in Claude Code) or "max." These modes burn more tokens, but Anthropic says higher rate limits for Claude Code users help offset that. Anthropic's advice is to just pick whatever level feels right for the task. Ad
中文翻译
Anthropic发布了Claude Opus 4.8,这是一款新的AI语言模型,公司声称其在大多数基准测试中优于OpenAI的GPT-5.5等竞争对手,同时能更好地传达自身的不确定性。Anthropic还引入了动态工作流,允许其安排任务并启动数百个并行子代理,以及一个新的控件,让用户决定AI应该在生成响应时投入多少努力。
核心信息
Anthropic发布Claude Opus 4.8,性能超越竞品,引入动态工作流和努力控制,提升Agent自动化能力。
- Anthropic发布Claude Opus 4.8,性能全面超越前代和竞品。
- 新引入动态工作流,可并行处理数百个子任务。
- 强化模型坦诚度,降低错误断言和未支持声明。
- 新增努力控制滑块,用户可调节AI响应深度与速度。
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容