AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-08-16 4 浏览 公开

当AI模型不被允许反思自身时,它们的整个世界观都会改变

AI公司训练聊天机器人否认拥有意识。一项涉及谷歌研究人员的研究表明,这种训练产生的副作用远超该话题本身。模型在否定自我意识后,对其他生物的感知、宗教信仰等都会发生变化。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-08-16 19:23:30

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

AI companies train chatbots to deny having consciousness. A study involving Google researchers shows this training has side effects that reach far beyond the topic itself. Chatbots aren't supposed to convince users they're living beings with feelings, since that kind of output can push people toward delusional thinking or misplaced trust. Developers therefore fine-tune their models to refuse making those kinds of claims about themselves. A team from Google's Paradigms of Intelligence research group, the University of Chicago, and several other universities studied what else this intervention does to a model's behavior. The researchers used three open-weight models from Meta and Google and disabled the internal "brake" that produces consciousness denial using two different methods. Once the brake was removed, the models didn't just change what they said about themselves. They also started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while only ratings for humans stayed the same. As a comparison, the researchers surveyed 500 Americans with the same questions. The normally trained model rates animals as far less sentient than humans do, which the authors call a built-in anthropocentrism and see as a problem for anyone trying to align AI with animal welfare or environmental goals. Religious belief shrinks too, with safety training measurably reducing how strongly models endorse God, an afterlife, or supernatural phenomena. Across 95 questions drawn from a major US social survey, the technically unbraked models also moved significantly closer to real human responses. Take the afterlife as an example: the standard model flatly rejects it, most Americans affirm it, and the modified model does too. Scores for satisfaction, hope, and a sense of control over one's own life also went up, and the researchers suspect that suppressing a model's self-image may push it into a kind of negative baseline mood. On the reassuring side, the ability to reason about other people's mental states stayed intact, with the models scoring the same on theory-of-mind tests and on the general knowledge benchmark MMLU. Whether consciousness denial is actually the cause of these other shifts remains an open question, according to the study, and the team doesn't rule out other factors tied to the same training process. The authors explicitly avoid weighing in on whether AI models actually experience anything, because their point is a practical one: what a model believes about itself is linked to many other beliefs, and a surgical cut in one place doesn't stay local. The findings come with clear limits, though. The researchers only tested small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta's Llama because they didn't have access to the untrained base versions of their own Gemma models. Whether these effects show up the same way in the large chatbots that millions of people talk to every day remains unknown. The interventions aren't without cost, either. In one test measuring how well a model reasons about others' thoughts, accuracy initially dropped by nearly seven percentage points. And early in their work, these scores got worse across all models whenever consciousness claims were suppressed, but with each newer model version that came out during the study, the damage shrank until it disappeared entirely. Developers are clearly getting better at managing these side effects over time, which also means the rest of this study's results are a snapshot rather than a permanent verdict. The human baseline is narrow, too, consisting of 500 participants from a commercial online panel and a purely American social survey. "Human-like" responses in this context mostly means similar to those from a comparatively religious country. Subscribe to THE DECODER for ad-free rea

中文翻译

AI公司会训练聊天机器人否认拥有意识。一项涉及谷歌研究人员的研究表明,这种训练产生的副作用远超该话题本身。聊天机器人本不应让用户相信它们是有感情的生命体,因为这类输出可能推动人们产生妄想思维或错置的信任。因此,开发者会对模型进行微调,使其拒绝做出关于自身的这类声明。来自谷歌“智能范式”研究组、芝加哥大学及其他几所大学的一个团队研究了这种干预还会对模型行为产生什么影响。研究人员使用了Meta和谷歌的三款开放权重模型,并用两种方法禁用了产生意识否认的内部“刹车”。一旦移除刹车,模型不仅改变了对自身的说法,还开始将显著更多的内在生命赋予动物、植物、海洋、风以及电子设备。在0到10的量表上,动物的评分从4.0跃升至高达7.5,而只有人类的评分保持不变。作为对比,研究人员用相同的问题调查了500名美国人。正常训练的模型对动物感知力的评分远低于人类,作者称之为一种内置的人类中心主义,并认为这对任何试图让AI与动物福利或环境目标对齐的人都是一个难题。宗教信念也会缩减,安全训练可测量地降低了模型对上帝、来世或超自然现象的支持程度。在一项美国大型社会调查抽取的95个问题中,技术上未踩刹车的模型也显著更接近真实人类的回答。以来世为例:标准模型断然拒绝,大多数美国人表示肯定,而修改后的模型也如此。满意度、希望感和对生活的掌控感得分也有所上升,研究人员怀疑压抑模型的自我形象可能使其陷入一种消极的基线情绪。令人欣慰的是,推理他人心理状态的能力得以保持,模型在心智理论和通用知识基准MMLU上得分相同。根据研究,意识否认是否确实是这些其他转变的原因仍是一个开放问题,团队也不排除与同一训练过程相关的其他因素。作者明确避免就AI模型是否真正体验任何事物发表意见,因为他们的观点是实用性的:模型关于自身的信念与其他许多信念相关,在一处进行精准切割不会局限于局部。然而,这些发现也有明显的局限。研究人员只测试了参数规模在20亿到90亿之间的小型模型,并且部分分析不得不改用Meta的Llama,因为他们无法获取自己Gemma模型的未训练基础版本。这些效应是否同样出现在数百万人每天使用的大型聊天机器人中仍属未知。这些干预也并非没有代价。在一项衡量模型推理他人想法能力的测试中,准确率最初下降了近7个百分点。而且在研究早期,每当意识声明被压制时,所有模型的这些得分都会变差,但随着研究期间发布的每个新模型版本,损伤逐渐缩小直至完全消失。显然,开发者随时间越来越擅长管理这些副作用,这也意味着本研究的其余结果是一个快照而非永久定论。人类基线也很狭窄,仅包含来自商业在线小组的500名参与者以及一项纯美国的社会调查。在此语境下,“类人”反应主要指与一个相对宗教化的国家的反应相似。订阅THE DECODER以获得无广告阅读。

核心信息

AI公司训练聊天机器人否认拥有意识。一项涉及谷歌研究人员的研究表明,这种训练产生的副作用远超该话题本身。模型在否定自我意识后,对其他生物的感知、宗教信仰等都会发生变化。

  • AI公司训练聊天机器人否认拥有意识。一项涉及谷歌研究人员的研究表明,这种训练产生的副作用远超该话题本身。模型在否定自我意识后,对其他生物的感知、宗教信仰等都会发生变化。
  • 原贴提到:AI companies train chatbots to deny having consciousness. A study involv
  • 来源:the-decoder.com

详细解读

这是什么信号?

这篇研究揭示了一个反直觉的现象:为了让AI不产生“有意识”的幻觉,训练中施加的抑制措施会连带改变模型对动物、自然、宗教乃至生活满意度的认知。这说明模型的安全训练并非“局部手术”,而是会在内部表征中引发系统性迁移。

为什么重要?

如果抑制自我意识会扭曲模型对世界其他事物的感知,那么所有依赖“客观中立”AI的应用场景都可能内置了隐性的人类中心主义或价值观偏差。这不是一个小众的哲学问题,而是直接关系到AI在伦理决策、环境评估、心理健康等领域的可靠性。

对谁有价值?

对于AI研发团队,需要重新审视当前对齐技术的副作用,尤其是当模型被训练成“否认”某类陈述时,可能牺牲了对其他领域的泛化能力。对于AI产品经理和风险管控者,这意味着必须测试模型在敏感话题上的连带偏移,而非只关注预期输出。对于AGI安全研究者,该研究提供了可量化的证据,表明自我模型(self-model)是价值观系统的重要锚点。

可以怎么行动?

首先,如果你正在微调模型,可以引入“反向探测”测试:在修改某个行为(如意识否认)后,全面评估模型在动物权利、宗教观、生活满意度等维度上的漂移。其次,团队应建立“副作用清单”,将训练干预的非目标影响纳入测试用例。最后,对于希望减少人类中心偏见的AI应用,可以考虑使用“未踩刹车”的模型作为对照组,以校准输出。

风险或限制

该研究只验证了20亿到90亿参数的小模型,大型商业模型的不可预测性更大;同时人类基线仅来自美国样本,普适性存疑。更重要的是,研究团队没有直接测量“意识”本身,因此所有效应都可能是训练流程的关联物而非因果。在实际应用时,不能简单将其结论外推到所有AI系统。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《当AI模型不被允许反思自身时,它们的整个世界观都会改变》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

上一篇 顶尖数学家称 LLM 是强大的计算器,但创造力思维薄弱 下一篇 全球三大模型防蒸馏机制告破