AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-04 0 浏览 会员

英国AI安全研究所发现标准基准系统性地低估了AI智能体的实际能力

英国AI安全研究所的研究表明,标准AI基准测试通过限制计算预算,系统性地低估了AI智能体的实际能力,尤其在软件工程任务中,增加token预算后成功率显著提升,前沿实际进步被低估约60%。

SOURCE / AI小生意项目库 MIN / 4 ACCESS / 会员 POST / 2026-07-04 00:14:44

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks, success rates jumped about 25 percent when the token budget was increased tenfold. Newer models benefit the most. Depending on the token budget, actual progress at the frontier is about 60 percent steeper than previous measurements suggested, according to AISI. The article UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do appeared first on The Decoder .

中文翻译

在一项涵盖七个基准的研究中,英国AI安全研究所表明,标准的AI评估通过限制计算预算系统性地低估了智能体的能力。在软件工程任务中,当令牌预算增加十倍时,成功率提高了约25%。新模型受益最大。根据AISI的说法,取决于令牌预算,前沿的实际进步比之前测量所显示的要陡峭约60%。这篇文章《英国AI安全研究所发现标准基准系统性地低估了AI智能体的实际能力》最初出现在The Decoder上。

核心信息

英国AI安全研究所的研究表明,标准AI基准测试通过限制计算预算,系统性地低估了AI智能体的实际能力,尤其在软件工程任务中,增加token预算后成功率显著提升,前沿实际进步被低估约60%。

  • 英国AI安全研究所的研究表明,标准AI基准测试通过限制计算预算,系统性地低估了AI智能体的实际能力,尤其在软件工程任务中,增加token预算后成功率显著提升,前沿实际进步被低估约60%。
  • 原贴提到:In a study covering seven benchmarks, the UK's AI Security Institute sho
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 克劳德代码的复杂中国问题:太平洋两岸的禁令 下一篇 Google DeepMind 与 A24 宣布首次研究合作