AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-04-13 0 浏览 会员

论文速读:DRBENCHER,聚焦形式化数学证明能力

本文介绍DRBENCHER,一个用于生成需要浏览和计算的合成问题基准生成器,覆盖生物化学、金融、地球物理、安全、历史五个领域,通过四个标准确保质量,人类评估有效率达76%,最强模型准确率仅20%。

SOURCE / AI技能杠杆 MIN / 4 ACCESS / 会员 POST / 2026-04-13 12:13:05

原贴

查看原文
作者:arXiv cs.AI 来源站点:arxiv.org 原贴时间:
论文速读:DRBENCHER,聚焦形式化数学证明能力

原文

arXiv:2604.09251v1 Announce Type: new Abstract: Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation, creating a blind spot in assessing real-world performance. We introduce DRBENCHER, a synthetic benchmark generator for questions that require both browsing and computation. It enforces four criteria: verifiability (gold answers are computed by executing parameterized code over knowledge-graph values), complexity (multi-hop entity identification, property retrieval, and domain-specific computation), difficulty (a two-stage verification cascade filters out questions solvable by the generating model), and diversity (a greedy max-min embedding filter maximizes coverage). These criteria are realized via a unified answer-first pipeline spanning five domains: biochemistry, financial, geophysical, security, and history. Human evaluation shows 76% validity (84% excluding stale data), with 35% of errors due to outdated knowledge-graph entries, highlighting an inherent limitation of systems that reason over evolving data. Automatic evaluation shows that the strongest frontier model achieves only 20% answer accuracy. Compared to manually constructed benchmarks (BrowseComp+, MATH-500, GPQA), DRBENCHER achieves the highest semantic diversity.

中文翻译

深度研究代理越来越多地将网页浏览与多步计算交织在一起,但现有基准孤立地评估这些能力,造成评估真实性能的盲点。我们提出DRBENCHER,一个针对需要浏览和计算的问题的合成基准生成器。它强制四个标准:可验证性(黄金答案通过执行知识图谱值上的参数化代码计算)、复杂性(多跳实体识别、属性检索和领域特定计算)、难度(两级验证级联过滤出生成模型可解的问题)和多样性(贪婪最大最小嵌入过滤器最大化覆盖)。这些标准通过统一的答案优先流程实现,涵盖五个领域:生物化学、金融、地球物理、安全和历史。人类评估显示76%的有效性(排除过时数据后为84%),其中35%的错误源于过时的知识图谱条目,凸显了推理演化数据系统的固有限制。自动评估显示最强前沿模型仅达到20%的答案准确率。与手动构建的基准(BrowseComp+、MATH-500、GPQA)相比,DRBENCHER实现了最高的语义多样性。

核心信息

本文介绍DRBENCHER,一个用于生成需要浏览和计算的合成问题基准生成器,覆盖生物化学、金融、地球物理、安全、历史五个领域,通过四个标准确保质量,人类评估有效率达76%,最强模型准确率仅20%。

  • DRBENCHER合成生成需浏览与计算的问题。
  • 覆盖生物化学等五个领域,强制四项标准。
  • 人类评估有效率达76%,35%错误因数据过时。
  • 最强模型准确率仅20%,远低于人类。
  • 语义多样性高于现有基准,更贴近真实场景。
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 论文速读:Camera Artist,聚焦 Agent 工作流自动化 下一篇 论文速读:SAGE,聚焦形式化数学证明能力