Knowledge File / AI技能杠杆
论文速读:DeepReviewer 2.0,解读最新 AI 进展
DeepReviewer 2.0是一种基于输出合约的自动化同行评审系统,通过可追溯的标注、局部证据和可执行行动生成审稿包,在ICLR 2025投稿上表现超越Gemini并优于人类审稿委员会。
SOURCE / AI技能杠杆
MIN / 4
ACCESS / 会员
POST / 2026-04-14 12:55:10
原贴
查看原文
原文
arXiv:2604.09590v1 Announce Type: new Abstract: Automated peer review is often framed as generating fluent critique, yet reviewers and area chairs need judgments they can \emph{audit}: where a concern applies, what evidence supports it, and what concrete follow-up is required. DeepReviewer~2.0 is a process-controlled agentic review system built around an output contract: it produces a \textbf{traceable review package} with anchored annotations, localized evidence, and executable follow-up actions, and it exports only after meeting minimum traceability and coverage budgets. Concretely, it first builds a manuscript-only claim--evidence--risk ledger and verification agenda, then performs agenda-driven retrieval and writes anchored critiques under an export gate. On 134 ICLR~2025 submissions under three fixed protocols, an \emph{un-finetuned 196B} model running DeepReviewer~2.0 outperforms Gemini-3.1-Pro-preview, improving strict major-issue coverage (37.26\% vs.\ 23.57\%) and winning 71.63\% of micro-averaged blind comparisons against a human review committee, while ranking first among automatic systems in our pool. We position DeepReviewer~2.0 as an assistive tool rather than a decision proxy, and note remaining gaps such as ethics-sensitive checks.
中文翻译
自动同行评审通常被描述为生成流畅的批评,但审稿人和领域主席需要的是可以审计的判断:问题在何处适用,什么证据支持它,以及需要什么具体的后续行动。DeepReviewer 2.0是一个过程可控的代理审稿系统,围绕输出合约构建:它生成一个可追溯的审稿包,包含锚定注释、局部证据和可执行的后续行动,并且仅在满足最低可追溯性和覆盖率预算后输出。具体来说,它首先构建一个仅基于手稿的声明-证据-风险台账和验证议程,然后执行议程驱动的检索,并在导出门限下编写锚定批评。在三个固定协议下的134篇ICLR 2025投稿上,运行DeepReviewer 2.0的未微调196B模型优于Gemini-3.1-Pro-preview,改进了严格的主要问题覆盖率(37.26%对23.57%),并在与人类审稿委员会的微观平均盲比较中赢得了71.63%的胜利,同时在我们的池中自动系统中排名第一。我们将DeepReviewer 2.0定位为辅助工具而非决策代理,并指出了剩余差距,如伦理敏感性检查。
核心信息
DeepReviewer 2.0是一种基于输出合约的自动化同行评审系统,通过可追溯的标注、局部证据和可执行行动生成审稿包,在ICLR 2025投稿上表现超越Gemini并优于人类审稿委员会。
- DeepReviewer 2.0通过输出合约生成可追溯审稿包,提升评审可信度。
- 在ICLR 2025测试中,其问题覆盖率超越Gemini,盲评优于人类委员会。
- 系统定位为辅助工具,需警惕伦理检查等缺口。
- 可帮助审稿人快速定位关键证据和具体行动建议。
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容