AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-04-20 0 浏览 会员

论文速读:DeepER-Med,评估 LLM Agent 表现

DeepER-Med是一个用于医学的深度证据研究框架,包含研究规划、代理协作和证据综合三个模块,配套100个专家级问题的数据集DeepER-MedQA,在多个指标上优于现有平台,并在8个临床案例中7个与临床建议一致。

SOURCE / AI技能杠杆 MIN / 9 ACCESS / 会员 POST / 2026-04-20 12:00:06

原贴

查看原文
作者:arXiv cs.AI 来源站点:arxiv.org 原贴时间:
论文速读:DeepER-Med,评估 LLM Agent 表现

原文

arXiv:2604.15456v1 Announce Type: new Abstract: Trustworthiness and transparency are essential for the clinical adoption of artificial intelligence (AI) in healthcare and biomedical research. Recent deep research systems aim to accelerate evidence-grounded scientific discovery by integrating AI agents with multi-hop information retrieval, reasoning, and synthesis. However, most existing systems lack explicit and inspectable criteria for evidence appraisal, creating a risk of compounding errors and making it difficult for researchers and clinicians to assess the reliability of their outputs. In parallel, current benchmarking approaches rarely evaluate performance on complex, real-world medical questions. Here, we introduce DeepER-Med, a Deep Evidence-based Research framework for Medicine with an agentic AI system. DeepER-Med frames deep medical research as an explicit and inspectable workflow of evidence-based generation, consisting of three modules: research planning, agentic collaboration, and evidence synthesis. To support realistic evaluation, we also present DeepER-MedQA, an evidence-grounded dataset comprising 100 expert-level research questions derived from authentic medical research scenarios and curated by a multidisciplinary panel of 11 biomedical experts. Expert manual evaluation demonstrates that DeepER-Med consistently outperforms widely used production-grade platforms across multiple criteria, including the generation of novel scientific insights. We further demonstrate the practical utility of DeepER-Med through eight real-world clinical cases. Human clinician assessment indicates that DeepER-Med's conclusions align with clinical recommendations in seven cases, highlighting its potential for medical research and decision support.

中文翻译

可信任性和透明性对于人工智能在医疗保健和生物医学研究中的临床采用至关重要。最近的深度研究系统旨在通过将AI代理与多跳信息检索、推理和综合相结合,加速基于证据的科学发现。然而,大多数现有系统缺乏明确且可检查的证据评估标准,造成错误累积的风险,并使研究人员和临床医生难以评估其输出的可靠性。同时,当前的基准测试方法很少评估在复杂的现实世界医学问题上的表现。在这里,我们介绍了DeepER-Med,一种用于医学的基于深度证据的研究框架,具有代理AI系统。DeepER-Med将深度医学研究构建为明确且可检查的基于证据生成的工作流程,由三个模块组成:研究规划、代理协作和证据综合。为了支持现实评估,我们还提出了DeepER-MedQA,一个基于证据的数据集,包含100个源自真实医学研究场景并由11位生物医学专家多学科小组策划的专家级研究问题。专家手动评估表明,DeepER-Med在多个标准上持续优于广泛使用的生产级平台,包括生成新颖的科学见解。我们进一步通过八个真实临床案例展示了DeepER-Med的实际效用。人类临床医生评估表明,DeepER-Med的结论在七个案例中与临床建议一致,突显了其在医学研究和决策支持方面的潜力。

核心信息

DeepER-Med是一个用于医学的深度证据研究框架,包含研究规划、代理协作和证据综合三个模块,配套100个专家级问题的数据集DeepER-MedQA,在多个指标上优于现有平台,并在8个临床案例中7个与临床建议一致。

  • DeepER-Med提出基于证据的医学研究AI框架。
  • 包含研究规划、代理协作、证据综合三模块。
  • 配套数据集DeepER-MedQA含100个专家级问题。
  • 实验表明DeepER-Med在多个指标上优于现有平台。
  • 在8个临床案例中7个与临床建议一致。
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 趋势解读:Introducing workspace agents in ChatGPT,聚焦 Agent 工作流自动化 下一篇 论文速读:Experience Compression Spectrum,解读最新研究结论