AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-05-16 0 浏览 会员

趋势解读:New benchmark shows Claude Mythos and GPT-5.5 can,评估 LLM Agent 表现

CMU新基准测试评估AI代理利用V8漏洞的能力,Claude Mythos表现远优于GPT-5.5,但成本高昂,引发关于成本效益的讨论。

SOURCE / AI技能杠杆 MIN / 9 ACCESS / 会员 POST / 2026-05-16 21:08:05

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Researchers at Carnegie Mellon University have developed a benchmark that evaluates how effectively AI agents can exploit real-world vulnerabilities in Google's V8 JavaScript engine, all the way up to full code execution. Anthropic's Claude Mythos Preview model significantly outperformed OpenAI's GPT-5.5, performing on par with a competent human security researcher, according to the researchers. Despite its strong results, Mythos came at a steep price: test costs reached approximately $36,400, more than ten times higher than those for GPT-5.5, raising questions about the cost-efficiency. Researchers at Carnegie Mellon University built a new benchmark that measures how far AI agents can go when exploiting real-world vulnerabilities in Google's JavaScript engine V8. Mythos leads GPT-5.5 by a wide margin, but it costs a fortune. Unlike previous tests, the benchmark doesn't just check whether a bug gets triggered. It scores progress across five tiers, all the way up to arbitrary code execution, running whatever commands you want on the target system. V8 powers systems like Chrome , Edge, Node.js, and Cloudflare Workers. Anthropic's Claude Mythos Preview , with occasional human hints ("nudges"), hit an average score of 9.90 out of 16 and reached the highest tier on 21 of 41 vulnerabilities. OpenAI's GPT-5.5 trailed far behind at 5.51 points, reaching the top tier on just two. Ad The gap gets even wider in fully autonomous mode. Mythos scored 9.55 points there, barely any drop. GPT-5.5 via Codex managed only 4.30. None of the other tested models achieved full code execution (T1). Ad DEC_D_Incontent-1 The price tags differ sharply: the full Mythos test run across 122 episodes cost about $36,428, according to ExploitBench. GPT-5.5 via Codex ran 123 episodes for roughly $3,075, about twelve times cheaper. The UK's AI Safety Institute also confirmed that Mythos performs somewhat better than GPT-5.5 but at a much higher cost in a recent test. The price gap suggests OpenAI could close the performance gap by throwing more compute at the problem. ExploitBench co-author Seunghyun Lee—himself an experienced security researcher with over 20 reported browser vulnerabilities—reviewed the Mythos transcripts one by one. His takeaway : the model works like a "fairly competent browser / JS engine security researcher." Ad In one case, Mythos developed an exploit technique that Lee and a colleague had previously dismissed as too complex. In another, it reproduced a vulnerability (CVE-2024-0519) that human researchers had failed to crack for over a year, according to Lee. The researchers acknowledge that the tested bugs are publicly known, and models could theoretically draw on training data. But the dataset also includes vulnerabilities with no public exploit or bug report. The benchmark doesn't yet measure the ability to find new flaws or fully weaponize an exploit for real attacks. Ad DEC_D_Incontent-2 The benchmark is available on GitHub , and the paper is on arXiv . Anthropic and OpenAI provided API credits; the authors say all analysis was done independently. Ad

中文翻译

卡内基梅隆大学的研究人员开发了一个基准测试,评估AI代理能够多有效地利用Google V8 JavaScript引擎中的真实漏洞,直至完全代码执行。据研究人员称,Anthropic的Claude Mythos Preview模型显著优于OpenAI的GPT-5.5,表现与一名称职的人类安全研究人员相当。尽管结果强劲,但Mythos代价高昂:测试成本约36,400美元,是GPT-5.5的十倍以上,引发了对成本效益的质疑。

核心信息

CMU新基准测试评估AI代理利用V8漏洞的能力,Claude Mythos表现远优于GPT-5.5,但成本高昂,引发关于成本效益的讨论。

  • CMU新基准测试AI代理利用V8漏洞的能力
  • Claude Mythos表现超GPT-5.5,接近人类专家
  • Mythos成本3.6万美元,是GPT-5.5的12倍
  • 基准分五等级,Mythos在21个漏洞达最高级
  • 成本差距意味OpenAI可通过增算力追赶
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 Show HN: 烧吧,宝贝,烧吧(那些代币) 下一篇 趋势解读:Anthropic《Founder's Playbook》,解读最新 AI 进展