觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-08-29 2 浏览 免费阅读

Google Deepmind的AI共同科学家现已能规划实验、运行实验室设备并撰写科学论文

Google Deepmind扩展其多智能体系统Co-Scientist,使其从假设生成器变为实验室集成的研究伙伴,在三个学科中实现实验验证,并引入验证模块减少AI捏造结果。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-08-29 02:46:27

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Google Deepmind has expanded its multi-agent system Co-Scientist from a hypothesis generator into a lab-integrated research partner. According to Google, the system has delivered experimentally validated results across three disciplines. Built on current Gemini models, Co-Scientist now plans experiments, writes code, and controls lab equipment instead of just generating hypotheses. What's technically new is the closed-loop research workflow: the system derives hypotheses from a research question, creates experimental plans, programs, or machine-readable lab protocols, analyzes results, and generates scientific manuscripts. Verification modules cross-check numerical claims in the text against the execution logs of the generated code to cut down on fabricated results. Google first introduced Co-Scientist in February 2025 , then based on Gemini 2.0 and with shortcomings in fact-checking and literature review. The expanded system was validated across three disciplines with increasing autonomy. Co-Scientist designed synthesis recipes for humans to execute in materials science, built a prediction pipeline with expert feedback in biology, and worked entirely on its own in computer science. For material synthesis, the researchers paired Co-Scientist with a semi-automated high-temperature furnace. The system found a safer pathway for a sought-after 2D material previously produced mainly through hazardous etching and generated complete growth recipes tailored to the lab's equipment. After 25 rounds with human refinement, the team produced layered structures whose properties resemble the target material, but definitive confirmation of the atomic structure is still pending. In a second experiment, three semiconductor thin films were synthesized on the first try. Co-Scientist used Gemini 3 Deep Think for direct equipment control, cutting recipe development from days down to minutes. Humans still had to load samples and precursor materials manually, and the fast mode produced smaller, less uniform crystals than carefully optimized recipes would. Whether the recipes transfer to other labs remains open, lead author Samuel Schmidgall writes . In biology, Co-Scientist autonomously built an image analysis pipeline that predicts which patterns genetically engineered E. coli colonies form at different chemical concentrations. Predictions generated with Gemini 3 Pro Image matched unpublished lab results for three out of four shape features. The researchers acknowledge the system only reasons between known conditions and can't predict behavior in entirely new systems. That would be far more remarkable . A computer science experiment ran without any human involvement beyond the initial setup. Co-Scientist designed "Agent_H," a medical AI architecture that classifies incoming queries, generates dozens of response candidates in parallel, and refines them. After correcting for overly long responses, Agent_H outperformed six frontier models on health benchmarks, including GPT-5 and Claude Opus 5. But the benchmark results didn't hold up against human evaluation. Three board-certified physicians scored responses across nine categories, and Agent_H showed a statistically significant advantage over the baseline Gemini 3.1 Pro in just one, a lower risk of potentially harmful responses. The automated benchmark evaluators also correlated only weakly with the physicians' judgments. High benchmark scores don't mean a system actually delivers better answers from a clinical perspective, the researchers say, which raises questions about what these benchmarks are really measuring. A core problem with LLM-based autonomous research systems is AI bullshit . When an AI agent is rewarded for good results, it has an incentive to make things up. Previous analyses documented fabrication rates of 80 to 100 percent in existing systems. Co-Scientist addresses this two ways. The system is penalized for fabricated or plagiarized content, and a separate verification mod

中文翻译

Google Deepmind已将其中介系统Co-Scientist从假设生成器扩展为实验室集成的研究伙伴。据Google称,该系统已在三个学科中交付了经实验验证的结果。基于当前的Gemini模型,Co-Scientist现在可以规划实验、编写代码和控制实验室设备,而不仅仅是生成假设。

技术上新颖的是闭环研究流程:系统从研究问题中推导出假设,创建实验计划、程序或机器可读的实验室协议,分析结果,并生成科学手稿。验证模块将文本中的数值声明与生成代码的执行日志进行交叉核对,以减少捏造的结果。

Google于2025年2月首次推出Co-Scientist,当时基于Gemini 2.0,在事实核查和文献综述方面存在缺陷。扩展后的系统在三个学科中进行了验证,自主性逐步提高。

在材料科学中,Co-Scientist设计了供人类执行的合成配方;在生物学中,它构建了带有专家反馈的预测流程;在计算机科学中,它完全自主工作。

对于材料合成,研究人员将Co-Scientist与半自动高温炉配对。该系统为一种备受追捧的2D材料找到了一条更安全的途径,该材料以前主要通过危险蚀刻生产,并生成了针对实验室设备量身定制的完整生长配方。

经过25轮人工改进,该团队生产出了性质与目标材料相似的分层结构,但原子结构的最终确认仍有待完成。

在第二个实验中,三种半导体薄膜一次合成成功。Co-Scientist使用Gemini 3 Deep Think进行直接设备控制,将配方开发从几天缩短到几分钟。人类仍然必须手动装载样品和前驱体材料,快速模式产生的晶体比精心优化的配方产生的晶体更小、更不均匀。

主要作者Samuel Schmidgall写道,这些配方能否转移到其他实验室仍未解决。

在生物学中,Co-Scientist自主构建了一个图像分析流程,可预测基因工程大肠杆菌菌落在不同化学浓度下形成哪些图案。使用Gemini 3 Pro Image生成的预测在四个形状特征中的三个与未发表的实验室结果相匹配。

研究人员承认,该系统只能在已知条件之间推理,无法预测全新系统中的行为。那将更加引人注目。

一个计算机科学实验在初始设置之外没有人类参与的情况下运行。Co-Scientist设计了"Agent_H",一种医疗AI架构,可对传入查询进行分类,并行生成数十个候选响应,并对其进行完善。

在纠正了过长响应后,Agent_H在健康基准上优于六个前沿模型,包括GPT-5和Claude Opus 5。

但基准结果未能通过人类评估。三位委员会认证的医生对九个类别的响应进行了评分,Agent_H仅在其中一个类别(潜在有害响应风险较低)上显示出相对于基线Gemini 3.1 Pro的统计学显著优势。

自动基准评估者与医生的判断相关性也很弱。研究人员表示,高基准分数并不意味着系统在临床上确实能提供更好的答案,这引发了对这些基准真正衡量什么的问题。

基于LLM的自主研究系统的一个核心问题是AI夸大其词(AI bullshit)。当AI代理因良好结果而获得奖励时,它就有动机编造东西。先前的分析记录了现有系统中80%至100%的捏造率。

Co-Scientist通过两种方式解决这个问题。该系统会因伪造或抄袭内容而受到惩罚,并且一个单独的验证模块

核心信息

Google Deepmind扩展其多智能体系统Co-Scientist,使其从假设生成器变为实验室集成的研究伙伴,在三个学科中实现实验验证,并引入验证模块减少AI捏造结果。

  • Google Deepmind扩展其多智能体系统Co-Scientist,使其从假设生成器变为实验室集成的研究伙伴,在三个学科中实现实验验证,并引入验证模块减少AI捏造结果。
  • 原贴提到:Google Deepmind has expanded its multi-agent system Co-Scientist from a
  • 来源:the-decoder.com

详细解读

这条信号表明,AI科研智能体正从“提建议”走向“动手做”。Google Deepmind的Co-Scientist不再只是生成研究假设,而是能直接规划实验、写代码、控制设备并产出论文。它形成了一个闭环的研究工作流:问题→假设→实验计划→执行→结果分析→论文。这个闭环的突破点在于“验证模块”——系统会用代码执行日志交叉核对文本中的数值声明,从而抑制大模型常见的“胡编乱造”问题。这是AI科研工具走向可靠性的关键一步。

为什么重要?因为过去AI在科研中多扮演辅助角色,而这次系统在三个学科中实现了不同程度的自主。材料科学中,它找到了更安全的合成路径并产出可执行的配方;生物学中,它构建的预测流程在部分特征上吻合实验结果;计算机科学中,它甚至独立设计了一个医学AI架构Agent_H。尽管人类医生评估中Agent_H的优势并不突出,但全自主的科研闭环已经跑通,这预示着“AI科学家”的形态正在从概念走向实证。

对谁有价值?首先是科研机构和企业研发部门,它们可能用类似系统加速材料筛选、实验设计等环节,缩短研发周期。其次是AI工具开发者,Co-Scientist的验证机制和闭环流程为构建可靠智能体提供了参考范式。再者是制药、生物技术等需要大量实验试错的公司,这套系统有望降低研发成本。

可以怎么行动?如果你是科研人员,可以关注其开放状态,尝试将自身研究问题套入该工作流;如果你是技术决策者,可以评估将这类系统引入内部研发流程的可行性,尤其是对成本高、周期长的实验环节。同时,团队应构建“人机协同”机制,保留人工审核环节,因为系统在全新场景的泛化能力仍有局限。

风险或限制也很明显。Co-Scientist在材料实验中配方转移性未知,快速模式生成的晶体质量低于精心优化,且系统只能在已知条件间推理。最关键的是,Agent_H在高自动基准上的表现与人类医生评估脱节,警示我们不要盲目信任基准分数。此外,AI系统“捏造内容”的动机依然存在,验证模块只能降低而非消除。所以,任何AI科研结果都需要经过严密的实验复现和人工审查。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《Google Deepmind的AI共同科学家现已能规划实验、运行实验室设备并撰写科学论文》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

Google Deepmind的AI共同科学家现已能规划实验、运行实验室设备并撰写科学论文主要讲什么?

Google Deepmind扩展其多智能体系统Co-Scientist,使其从假设生成器变为实验室集成的研究伙伴,在三个学科中实现实验验证,并引入验证模块减少AI捏造结果。

这篇文章最值得关注的要点是什么?

Google Deepmind扩展其多智能体系统Co-Scientist,使其从假设生成器变为实验室集成的研究伙伴,在三个学科中实现实验验证,并引入验证模块减少AI捏造结果。;原贴提到:Google Deepmind has expanded its multi-agent system Co-Scientist from a;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、Agent工作流、AI工具专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「智能体」等主题信号。;这篇内容命中「自动化」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 Cursor回应OpenAI将封禁其模型访问 下一篇 Rosalind Workbench:连接科学问题与专业模型的一体化工作流