AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-23 0 浏览 会员

SymptomAI:面向日常症状评估的对话式AI代理

谷歌研究团队开展了一项全国规模研究,测试对话式AI代理SymptomAI在症状评估和鉴别诊断中的表现。通过与真实临床诊断对比及Fitbit生物信号分析,显示AI在传染病诊断方面与生理信号吻合,验证了其有效性。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-23 05:32:00

原贴

查看原文
作者:Google Research Blog 来源站点:research.google 原贴时间:

原文

Joseph Breda, Student Researcher, and Jake Sunshine, Research Scientist, Google Research We present a first-of-its-kind research of AI for differential diagnosis and symptom checking through a national-scale study. A large proportion of clinical diagnoses can be derived from language-based interviews alone. These diagnostic interviews are typically conducted by clinicians through doctor-patient interactions during in-person or remote visits. While these interactions are the gold standard for symptom assessment, they can often suffer from financial , geographic , and systemic barriers that limit their accessibility. Current language models (LMs) have demonstrated strong differential diagnosis assessment capabilities when evaluated on curated medical case-studies, highlighting their potential to support the diagnostic process. However, existing evaluations have largely relied on curated, highly detailed and sometimes synthetic patient vignettes, which may not reflect real world experience and clinical presentation variability. These evaluations do not capture how everyday patients report their health symptoms, for example with varying levels of medical literacy, incomplete information, and other complexities that arise through natural conversation. This represents a key gap, leading to uncertainty of how LMs might perform in real-world contexts. To address this gap, we conduct an in-situ comparative research study of a set of experimental conversational prototype AI agents designed to explore how conversational AI might conduct end-to-end symptom interviews and differential diagnostic assessment for research benchmarking purposes. In our recent research paper, “ SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment ”, we share results from a randomized national scale study (n=13,917) in which consented research participants interact with one of five possible Gemini Flash 2.0 SymptomAI agents. All diagnoses, labels, and disease associations generated during the study were for research analysis only and did not constitute confirmed clinical diagnoses or official medical assessments. Two weeks after their interaction with the AI agents, we asked research participants to report any diagnoses they received from a visit with a healthcare provider. Using this data, we conducted a clinical expert annotation study comparing SymptomAI’s diagnostic performance relative to real clinicians' medical assessments. After assessing the accuracy of SymptomAI’s differential diagnoses (DDx), we further compare SymptomAI’s diagnoses against biosignals from participants’ Fitbit wearable devices in the time leading up to their conversation with SymptomAI. We show that SymptomAI conversations that led to diagnosis with an infectious disease etiology coincide with physiological trends that may indicate an immune response, suggesting further evidence of SymptomAI’s performance. For this research, we administered a set of conversational AI agents for end-to-end patient interviewing and differential diagnostic assessment. A participant could converse with an agent about their symptoms, receive a candidate differential list and enter a subsequent diagnosis. We enrolled 13,917 consenting research study participants who each describe their symptoms to one of five randomized SymptomAI agents, each with varying degrees of flexibility in how they conducted the symptom interview. During these conversations, participants described their symptoms and SymptomAI asked follow-up questions, with conversations culminating in a final differential diagnosis (DDx, a list of plausible diagnoses) and recommendations for next steps. Participants could then go on to see a healthcare provider and were asked to share the outcome of that visit via a survey two-weeks later. To evaluate and baseline SymptomAI’s assessment, we conducted a clinical-expert annotation study in which a panel of three board-certified clinicians reviewed the conversation transc

中文翻译

Joseph Breda,学生研究员,以及Jake Sunshine,研究科学家,谷歌研究院。我们首次提出了一项通过全国规模研究进行的AI鉴别诊断和症状检查研究。很大一部分临床诊断仅凭基于语言的访谈即可得出。这些诊断性访谈通常由临床医生通过面对面或远程就诊时的医患互动进行。虽然这些互动是症状评估的金标准,但它们常常受到经济、地理和系统性障碍的限制,影响了可及性。当前的语言模型在基于精选医学案例研究进行评估时,已显示出强大的鉴别诊断能力,凸显了它们支持诊断过程的潜力。然而,现有的评估主要依赖于精心挑选、高度详细甚至有时是合成的患者小传,这些可能无法反映现实世界的经验和临床表现的变异性。这些评估没有捕捉到日常患者如何报告他们的健康症状,例如不同程度的医学素养、不完整的信息以及自然对话中出现的其他复杂性。这代表了一个关键差距,导致语言模型在现实世界中可能表现如何的不确定性。为了解决这一差距,我们进行了一项对一组实验性对话原型AI代理的现场比较研究,旨在探索对话式AI如何为研究基准目的进行端到端的症状访谈和鉴别诊断评估。在我们最近的研究论文《SymptomAI:面向日常症状评估的对话式AI代理》中,我们分享了一项随机全国规模研究(n=13,917)的结果,其中同意的研究参与者与五种可能的Gemini Flash 2.0 SymptomAI代理之一进行互动。研究期间生成的所有诊断、标签和疾病关联仅用于研究分析,并不构成确认的临床诊断或官方医学评估。在与AI代理互动两周后,我们要求研究参与者报告他们从医疗服务提供者就诊中收到的任何诊断。利用这些数据,我们进行了一项临床专家注释研究,比较SymptomAI的诊断性能与真实临床医生的医学评估。在评估了SymptomAI的鉴别诊断(DDx)准确性之后,我们进一步将SymptomAI的诊断与参与者与SymptomAI对话前来自Fitbit可穿戴设备的生物信号进行比较。我们表明,导致感染性疾病病因诊断的SymptomAI对话与可能指示免疫反应的生理趋势一致,这进一步证明了SymptomAI的性能。在这项研究中,我们实施了一组端到端患者访谈和鉴别诊断评估的对话AI代理。参与者可以与代理讨论他们的症状,接收一个候选鉴别列表并进入后续诊断。我们招募了13,917名同意的研究参与者,每人向五个随机化SymptomAI代理之一描述其症状,这些代理在如何进行症状访谈方面具有不同程度的灵活性。在这些对话中,参与者描述他们的症状,SymptomAI提出后续问题,对话最终给出一个最终的鉴别诊断(DDx,一个可能的诊断列表)和下一步建议。参与者随后可以去看医疗服务提供者,并在两周后通过调查分享就诊结果。为了评估和建立SymptomAI评估的基线,我们进行了一项临床专家注释研究,由三位委员会认证的临床医生组成的小组审查了对话记录。

核心信息

谷歌研究团队开展了一项全国规模研究,测试对话式AI代理SymptomAI在症状评估和鉴别诊断中的表现。通过与真实临床诊断对比及Fitbit生物信号分析,显示AI在传染病诊断方面与生理信号吻合,验证了其有效性。

  • 谷歌研究团队开展了一项全国规模研究,测试对话式AI代理SymptomAI在症状评估和鉴别诊断中的表现。通过与真实临床诊断对比及Fitbit生物信号分析,显示AI在传染病诊断方面与生理信号吻合,验证了其有效性。
  • 原贴提到:Joseph Breda, Student Researcher, and Jake Sunshine, Research Scientist,
  • 来源:research.google
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 TheNumbers.com因AI爬虫与安全攻击导致网站崩溃重建 下一篇 Alphabet Q2:AI 投资推动营收增长 24%,Gemini 月活达 9.5 亿