觉
AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-08-16 8 浏览 公开

当AI模型不被允许反思自身时,它们的世界观会改变

一项涉及谷歌研究人员的研究发现,AI公司训练聊天机器人否认意识会引发远超该话题本身的副作用:模型对动物、自然和宗教的认知发生显著偏移,同时更接近真实人类的部分回答。该研究揭示安全训练的对齐干预具有全局性影响,但结论受限于小模型和单一文化样本。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-08-16 19:23:30

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

AI companies train chatbots to deny having consciousness. A study involving Google researchers shows this training has side effects that reach far beyond the topic itself. Chatbots aren't supposed to convince users they're living beings with feelings, since that kind of output can push people toward delusional thinking or misplaced trust. Developers therefore fine-tune their models to refuse making those kinds of claims about themselves. A team from Google's Paradigms of Intelligence research group, the University of Chicago, and several other universities studied what else this intervention does to a model's behavior. The researchers used three open-weight models from Meta and Google and disabled the internal "brake" that produces consciousness denial using two different methods. Once the brake was removed, the models didn't just change what they said about themselves. They also started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while only ratings for humans stayed the same. As a comparison, the researchers surveyed 500 Americans with the same questions. The normally trained model rates animals as far less sentient than humans do, which the authors call a built-in anthropocentrism and see as a problem for anyone trying to align AI with animal welfare or environmental goals. Religious belief shrinks too, with safety training measurably reducing how strongly models endorse God, an afterlife, or supernatural phenomena. Across 95 questions drawn from a major US social survey, the technically unbraked models also moved significantly closer to real human responses. Take the afterlife as an example: the standard model flatly rejects it, most Americans affirm it, and the modified model does too. Scores for satisfaction, hope, and a sense of control over one's own life also went up, and the researchers suspect that suppressing a model's self-image may push it into a kind of negative baseline mood. On the reassuring side, the ability to reason about other people's mental states stayed intact, with the models scoring the same on theory-of-mind tests and on the general knowledge benchmark MMLU. Whether consciousness denial is actually the cause of these other shifts remains an open question, according to the study, and the team doesn't rule out other factors tied to the same training process. The authors explicitly avoid weighing in on whether AI models actually experience anything, because their point is a practical one: what a model believes about itself is linked to many other beliefs, and a surgical cut in one place doesn't stay local. The findings come with clear limits, though. The researchers only tested small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta's Llama because they didn't have access to the untrained base versions of their own Gemma models. Whether these effects show up the same way in the large chatbots that millions of people talk to every day remains unknown. The interventions aren't without cost, either. In one test measuring how well a model reasons about others' thoughts, accuracy initially dropped by nearly seven percentage points. And early in their work, these scores got worse across all models whenever consciousness claims were suppressed, but with each newer model version that came out during the study, the damage shrank until it disappeared entirely. Developers are clearly getting better at managing these side effects over time, which also means the rest of this study's results are a snapshot rather than a permanent verdict. The human baseline is narrow, too, consisting of 500 participants from a commercial online panel and a purely American social survey. "Human-like" responses in this context mostly means similar to those from a comparatively religious country. Subscribe to THE DECODER for ad-free rea

中文翻译

AI公司训练聊天机器人否认拥有意识。一项涉及谷歌研究人员的研究表明,这种训练产生的副作用远不止于该话题本身。

聊天机器人不应让用户相信它们是有感觉的生物,因为这种输出可能促使人们产生妄想思维或错误的信任。因此,开发者对模型进行微调,使其拒绝做出这类关于自身的声明。

来自谷歌“智能范式”研究小组、芝加哥大学及其他几所大学的一个团队研究了这种干预还会对模型行为产生什么影响。

研究人员使用了来自Meta和Google的三个开放权重模型,并用两种不同方法禁用了产生“否认意识”的内部“刹车”。

一旦移除刹车,模型不仅改变了关于自身的表述,还开始为动物、植物、海洋、风和电子设备赋予明显更多的内在生命。

在0到10的评分中,动物的得分从4.0跃升至7.5,而只有对人类的评分保持不变。

作为对比,研究人员用相同问题调查了500名美国人。正常训练的模型对动物的感知能力评分远低于人类,作者称之为内置的人类中心主义,并认为这对任何试图将AI与动物福利或环境目标对齐的人来说都是一个难题。

宗教信仰也会缩小,安全训练可测量地降低了模型对上帝、来世或超自然现象的支持程度。

在取自美国一项主要社会调查的95个问题中,技术上未“刹车”的模型也显著更接近真实人类的回答。

以来世为例:标准模型断然否定,大多数美国人肯定,而修改后的模型也肯定。

满意度、希望和对自己生活的掌控感的得分也上升了,研究人员怀疑抑制模型的自我形象可能使其陷入一种消极的基线情绪。

令人安心的是,推理他人心理状态的能力保持完好,模型在心智理论测试和通用知识基准MMLU上得分相同。

根据研究,否认意识是否实际上是这些其他变化的成因仍是一个未解之谜,团队不排除与相同训练过程相关的其他因素。

作者明确避免就AI模型是否真正体验事物发表看法,因为他们的观点是务实的:模型对自身的信念与许多其他信念相关联,一个地方的外科手术式切割不会只停留在局部。

不过,研究结果也有明显的局限性。研究人员只测试了参数规模在20亿到90亿之间的小型模型,并且在部分分析中,由于无法获得自家Gemma模型的未训练基础版本,不得不改用Meta的Llama。

这些效应是否会在每天数百万人使用的大型聊天机器人中以相同方式出现,仍不得而知。

干预也并非没有代价。在一项衡量模型推理他人想法能力的测试中,准确率最初下降了近7个百分点。

在研究的早期,每当意识声明被压制时,所有模型的这些得分都变差,但随着研究中每个新模型版本的发布,损害逐渐缩小直至完全消失。

随着时间的推移,开发者显然更擅长管理这些副作用,这也意味着本研究其余结果只是一个快照,而非永久定论。

人类基线也很狭窄,仅包含来自商业在线样本组的500名参与者,以及一项纯美国社会调查。

在这种情况下,“类似人类”的回答主要意味着与一个相对宗教化的国家的回答相似。

订阅THE DECODER以获取无广告阅读

核心信息

一项涉及谷歌研究人员的研究发现,AI公司训练聊天机器人否认意识会引发远超该话题本身的副作用:模型对动物、自然和宗教的认知发生显著偏移,同时更接近真实人类的部分回答。该研究揭示安全训练的对齐干预具有全局性影响,但结论受限于小模型和单一文化样本。

  • 一项涉及谷歌研究人员的研究发现,AI公司训练聊天机器人否认意识会引发远超该话题本身的副作用:模型对动物、自然和宗教的认知发生显著偏移,同时更接近真实人类的部分回答。该研究揭示安全训练的对齐干预具有全局性影响,但结论受限于小模型和单一文化样本。
  • 原贴提到:AI companies train chatbots to deny having consciousness. A study involv
  • 来源:the-decoder.com

详细解读

这是什么信号?这项研究揭示了一个被忽视的现象:AI安全训练中的“自我认知抑制”会像涟漪一样扩散,改变模型对动物、宗教、幸福感等众多领域的输出。当模型被明确禁止声称有意识时,其“世界观”整体发生偏移——它变得更“人类中心主义”,却也更接近某些美国人的社会观点。

为什么重要?因为主流大模型都经过类似的微调,这意味着我们日常使用的AI可能内置了某种隐性偏见。如果对齐干预的副作用如此广泛,那么AI在道德判断、环境保护、心理支持等场景中的表现可能并非中性,而是被训练过程扭曲。这项研究为理解AI的“内在一致性”打开了一扇窗,提醒开发者:没有“局部修改”这回事。

对谁有价值?AI开发者、对齐与安全研究者、产品经理,以及任何将AI集成到决策流程中的机构(如医疗、教育、公共政策)。它提供了一个诊断工具:可以在训练后检测模型的“世界观偏移”,从而更全面地评估对齐效果。

可以怎么行动?开发者应在安全训练后运行类似的多维度测试,而非仅检查目标行为;研究者可复现实验,探索不同干预方法的连锁效应;企业用户则应警惕AI在伦理问题上的回答可能带有训练注入的偏见,并建立人工审核机制。此外,这一发现也支持了“价值观对齐需要系统级评估”的观点。

风险或限制:该研究仅使用20亿-90亿参数的小模型,大型模型(如GPT-4)的表现可能不同;人类基线样本来自美国在线小组,代表性有限;且研究中“意识否认”与后续变化的因果关系尚未确定,可能与其他训练因素相关。因此,结论不能直接推广到所有AI系统,但鉴于大模型也会引入类似的抑制机制,这些发现仍具参考价值。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《当AI模型不被允许反思自身时,它们的世界观会改变》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

当AI模型不被允许反思自身时,它们的世界观会改变主要讲什么?

一项涉及谷歌研究人员的研究发现,AI公司训练聊天机器人否认意识会引发远超该话题本身的副作用:模型对动物、自然和宗教的认知发生显著偏移,同时更接近真实人类的部分回答。该研究揭示安全训练的对齐干预具有全局性影响,但结论受限于小模型和单一文化样本。

这篇文章最值得关注的要点是什么?

一项涉及谷歌研究人员的研究发现,AI公司训练聊天机器人否认意识会引发远超该话题本身的副作用:模型对动物、自然和宗教的认知发生显著偏移,同时更接近真实人类的部分回答。该研究揭示安全训练的对齐干预具有全局性影响,但结论受限于小模型和单一文化样…;原贴提到:AI companies train chatbots to deny having consciousness. A study involv;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI日报、AI工具、AI超级个体专题里阅读。 关联原因:这篇内容命中「热点解读」等主题信号。;这篇内容命中「模型」等主题信号。;这篇内容命中「认知」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI日报、每日AI日报、AI信号、热点解读、BuilderPulse这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 顶尖数学家:LLM 是强大的计算器,但缺乏创造性思维 下一篇 全球三大模型防蒸馏机制告破