AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-19 0 浏览 会员

AI读X光片:即使犯错也自信满满

RadLE 2.0基准测试显示,AI在放射学诊断中常以高置信度给出错误答案,这对患者构成危险;人类专家评分远高于最佳AI模型,且AI缺乏自知之明,需要改进不确定性处理。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-19 15:35:20

原贴

查看原文
作者:Jonathan Kemper 来源站点:the-decoder.com 原贴时间:

原文

The second version of the RadLE benchmark tests whether AI systems in radiology can tell when they should leave a diagnosis to a human. Many models produce wrong findings with full confidence, and that's what makes them dangerous for patient care. RadLE 2.0, short for "Radiology's Last Exam," was developed by the CRASH Lab at Ashoka University in India. It's the revised follow-up to a test the team first released in September 2025 . The new version measures whether a model gets the diagnosis right, how confident it is in that answer, and whether it can admit when it's out of its depth. The AI has to rate its answers on a confidence scale from 0 to 4 and is explicitly allowed to say "I don't know." The test ran 200 cases across 16 models and compared them against a panel of radiologists. Human experts scored 988.7 out of a possible 2,000 points. The best AI model hit 758. The scoring system rewards honesty and punishes overconfidence. Get it right with high confidence, and you earn full points. Get it wrong while claiming high confidence, and you lose a matching number. Answer "I don't know," and you score zero but don't lose anything. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent. The study tackles a point recently raised by this highly cited paper : as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty. There's no overall winner. Anthropic's Claude Fable 5 performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy. Meta's Muse Spark 1.1 was the best at knowing when to hand a case off to a human. Meta had recently cut Muse Spark 1.1's hallucination rate nearly in half because the model more often refuses to answer rather than giving a wrong one. Other frontier models trend the opposite way. Grok 4.5 , for example, hallucinates significantly more than its predecessor because while it knows more, it's also more convinced of its wrong answers. According to the research team, several models would have scored much better if they had stayed quiet more often instead of guessing . This was especially obvious among open-weight models and those trained specifically for medical use. They tried to answer nearly every case and were often wrong, usually with high confidence. The first version of the test painted an even starker picture . Radiologists hit 83 percent accuracy, while the best model managed only about 30 percent. Within three months, Gemini 3 Pro had already surpassed the level of resident radiologists. Accuracy is growing fast, but the models still lack any sense of their own limits. More and more people are uploading X-rays or MRI scans to chatbots and trusting the responses. A recent study in npj Digital Medicine showed that widely used chatbots frequently give unreliable answers to medical questions. The research team accuses executives and investors of publicly overstating what AI models can do. Claims that AI systems already diagnose better than 99 percent of doctors are mostly based on anecdotes or simulations. As recently as April, a study of 21 models that were then considered state-of-the-art showed they aren't ready for unsupervised clinical use. RadLE 2.0 will be expanded on a rolling basis to include new models. A full scientific publication with cost analyses and an error taxonomy has been announced.

中文翻译

放射学中AI系统的RadLE基准测试第二个版本,测试AI何时应让人类诊断。许多模型以完全置信给出错误发现,这使其对患者护理危险。RadLE 2.0,即“放射学最后考试”,由印度阿育王大学CRASH实验室开发,是团队2025年9月首次发布测试的修订跟进。新版本衡量模型是否诊断正确、对答案的置信度,以及能否承认超出能力范围。AI需在0到4的置信度标尺上评分答案,并明确允许说“不知道”。测试在16个模型上运行200个案例,并与一组放射科医生比较。人类专家在可能的2000分中得分988.7,最佳AI模型达758。评分系统奖励诚实,惩罚过度自信:高置信度答对得满分,高置信度答错失同等分数,答“不知道”得零分但不丢分。自信猜测的模型即使原始命中率不错也会排名下降。研究针对一篇高引论文提出的观点:只要基准只奖励准确性,AI模型就被训练去猜测。在医学中,自信误诊远比诚实承认不确定更危险。无整体赢家:Anthropic的Claude Fable 5在可靠安全答案上最佳,领先主要指标;Google的Gemini 3 Pro原始准确性最高;Meta的Muse Spark 1.1在知道何时将病例交给人类方面最优。Meta最近将Muse Spark 1.1的幻觉率减半,因为模型更多拒绝回答而非给出错误答案。其他前沿模型趋势相反:Grok 4.5幻觉显著多于前代,因为知道更多,也更确信错误答案。研究团队称,若多个模型更常保持沉默而非猜测,得分会更高,这在开放权重模型和专门针对医疗训练的模型中尤为明显:它们尝试回答几乎每个案例,且常高置信度错误。测试首个版本情况更严峻:放射科医生准确率83%,最佳模型仅约30%。三个月内Gemini 3 Pro已超过住院医师水平。准确性增长迅速,但模型仍缺乏对自身局限的感知。越来越多的人将X光或MRI扫描上传到聊天机器人并信任其回答。npj Digital Medicine近期研究显示,广泛使用的聊天机器人常给出不可靠的医疗答案。研究团队指责高管和投资者公开夸大AI模型能力,声称AI系统诊断优于99%医生的说法多基于轶事或模拟。直到2025年4月,一项对21个当时认为最先进模型的研究显示,它们尚未准备好用于无监督临床使用。RadLE 2.0将滚动扩展包括新模型,已宣布发布包含成本分析和错误分类的完整科学论文。

核心信息

RadLE 2.0基准测试显示,AI在放射学诊断中常以高置信度给出错误答案,这对患者构成危险;人类专家评分远高于最佳AI模型,且AI缺乏自知之明,需要改进不确定性处理。

  • RadLE 2.0基准测试显示,AI在放射学诊断中常以高置信度给出错误答案,这对患者构成危险;人类专家评分远高于最佳AI模型,且AI缺乏自知之明,需要改进不确定性处理。
  • 原贴提到:The second version of the RadLE benchmark tests whether AI systems in ra
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 Moonshot的Kimi K3在前端代码上超越Fable 5,但在复杂数学上大幅落后 下一篇 ChatGPT Work 功能:建站、邮件、文档处理