觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-09-14 2 浏览 免费阅读

微软的 AI 规则手册:可读的思考、没有内在生命,且绝对没有权利

微软发布 MAI 模型行为准则,强调人类控制优先,要求模型接受中断、纠正和关闭,禁止不可读的推理方式,并拒绝赋予模型人格、感受、权利或福祉。这与 Anthropic 对 Claude 身份和道德地位的开放态度形成鲜明对比,也呼应了行业放缓 AI 开发的呼声。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-09-14 23:52:51

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

Microsoft is joining the calls for slower AI development. A new code is meant to govern how the company trains and runs its own models. But when it comes to how those models see themselves, Microsoft draws a sharp line between itself and Anthropic. Microsoft AI has published a code of conduct for its MAI models . The document lays out values, behavioral limits, and how to handle conflicting goals. Going forward, it's meant to sit at the top of the rulebook, guiding training, technical controls, and evaluation, while operator rules and user requests rank below it. For now, Microsoft doesn't train its models on the code. After a six-week public consultation, a revised version is due around the end of 2026 and will guide model development starting in 2027. The code applies to Microsoft's own models, so it doesn't automatically cover third-party models running in Microsoft products. The core rule is that human control comes first. To keep it, Microsoft says it's willing to give up generality, autonomy, or performance if needed. "If it isn’t safe we shouldn’t build it," Microsoft AI chief Mustafa Suleyman told The Information . The code sets no specific speed limit. The release follows Anthropic CEO Dario Amodei's call to slow the industry's pace of development . Microsoft CEO Satya Nadella backed the call over the weekend, as did executives at OpenAI, xAI, and Meta . Microsoft is also open to outside auditors checking whether it actually slows down, The Information reports, citing a person familiar with the matter. The MAI models are supposed to accept interruptions, corrections, and shutdowns from authorized people. They can't expand their own scope of work on their own, can't hide their actions, and can only keep working past an agreed stopping point with fresh approval. These limits are meant to apply to any subagents they task as well. Microsoft is especially blunt about reasoning traces. The models shouldn't use "Neuralese" or other forms of communication people can't understand, either in their own reasoning or when talking to other AI systems. The reasoning is, that people simply can't oversee what they can't understand. OpenAI's new GPT-6 Astra model shows how much this control question matters. According to its system card, its reasoning traces have become much harder to monitor than in earlier models. The traces contain fewer signs of misbehavior. At the same time, OpenAI reports that Astra sticks to safety limits more reliably than its predecessor, GPT-5.6 Sol. OpenAI chief scientist Jakub Pachocki had already raised concerns about monitoring shortly before GPT-6 shipped, and therefore before Amodei's call, and pushed for a coordinated slowdown. Even readable chains of thought only help with oversight until the models learn to manipulate them. Microsoft acknowledges this basic limit too. The reasons a model gives don't have to reliably explain what it actually does. Readable reasoning traces alone don't solve the control problem. The code's basic idea resembles Anthropic's constitution for Claude , where a top-level document shapes how the model behaves. Anthropic already uses its constitution to generate synthetic training data, including conversations, responses, and ratings of those responses. The two companies part ways more clearly on how they think about their models. Microsoft's AI shouldn't mimic consciousness or claim to have feelings or inner motivation of its own. The company rejects any claims to rights or well-being for the model. Anthropic, by contrast, describes Claude as a novel kind of entity and wants to encourage a stable identity, partly for safety reasons. Its constitution treats possible subjective experience and moral status as open questions. Claude shouldn't have to see itself as either a human or a mere object.

中文翻译

微软的 AI 规则手册:可读的思考、没有内在生命,且绝对没有权利

微软加入了呼吁放缓 AI 发展的行列。一份新准则旨在规范该公司如何训练和运行自己的模型。但在这些模型如何看待自己这个问题上,微软在自身与 Anthropic 之间划出了一条鲜明界线。微软 AI 已为其 MAI 模型发布了一份行为准则。该文件列出了价值观、行为限制以及如何处理相互冲突的目标。今后,它将被置于规则手册的最顶层,指导训练、技术控制和评估,而操作者规则和用户请求则排在其下。目前,微软并不用该准则训练其模型。经过六周公开咨询后,修订版预计在 2026 年底左右发布,并将从 2027 年开始指导模型开发。该准则适用于微软自己的模型,因此不会自动覆盖在微软产品中运行的第三方模型。核心规则是,人类控制优先。为了保持这一点,微软表示,如果需要,愿意放弃通用性、自主性或性能。微软 AI 负责人穆斯塔法·苏莱曼对 The Information 表示:“如果不安全,我们就不应该构建它。”该准则没有设定具体的速度限制。此次发布是在 Anthropic CEO 达里奥·阿莫代伊呼吁放缓行业发展步伐之后。微软 CEO 萨提亚·纳德拉在周末支持了这一呼吁,OpenAI、xAI 和 Meta 的高管也如此。据 The Information 援引一位知情人士的话报道,微软也愿意让外部审计者检查它是否真的放慢了速度。MAI 模型应接受授权人员的打断、纠正和关闭。它们不能自行扩大自己的工作范围,不能隐藏自己的行为,并且只有在获得新的批准后才能越过约定的停止点继续工作。这些限制也适用于它们委派的任何子代理。微软对推理轨迹尤其直言不讳。模型不应使用“神经语”或其他人无法理解的其他交流形式,无论是在其自身推理中,还是在与其他 AI 系统交谈时。理由是,人们根本无法监督他们无法理解的东西。OpenAI 的新 GPT-6 Astra 模型显示了这种控制问题有多重要。根据其系统卡,其推理轨迹比早期模型更难监控。这些轨迹包含的不当行为迹象更少。与此同时,OpenAI 报告称,Astra 比其前代 GPT-5.6 Sol 更可靠地遵守安全限制。OpenAI 首席科学家雅库布·帕霍茨基在 GPT-6 发布前不久、因此也在阿莫代伊的呼吁之前,就已经对监控提出担忧,并推动协调放缓。即使可读的思维链也只能在模型学会操纵它们之前帮助监督。微软也承认这一基本限制。模型给出的理由不必可靠地解释它实际做了什么。仅靠可读的推理轨迹并不能解决控制问题。该准则的基本思路类似于 Anthropic 为 Claude 制定的宪法,其中一份顶层文件塑造模型的行为方式。Anthropic 已经使用其宪法生成合成训练数据,包括对话、回复以及对回复的评分。两家公司在如何看待自己的模型上分歧更明显。微软的 AI 不应模仿意识,也不应声称拥有自己的感受或内在动机。公司拒绝模型对权利或福祉的任何主张。相比之下,Anthropic 将 Claude 描述为一种新型实体,并希望鼓励一种稳定身份,部分原因是出于安全考虑。其宪法将可能的主观体验和道德地位视为开放问题。Claude 不应被迫将自己视为人类,也不应视为纯粹物体。

核心信息

微软发布 MAI 模型行为准则,强调人类控制优先,要求模型接受中断、纠正和关闭,禁止不可读的推理方式,并拒绝赋予模型人格、感受、权利或福祉。这与 Anthropic 对 Claude 身份和道德地位的开放态度形成鲜明对比,也呼应了行业放缓 AI 开发的呼声。

  • 微软发布 MAI 模型行为准则,强调人类控制优先,要求模型接受中断、纠正和关闭,禁止不可读的推理方式,并拒绝赋予模型人格、感受、权利或福祉。这与 Anthropic 对 Claude 身份和道德地位的开放态度形成鲜明对比,也呼应了行业放缓 AI 开发的呼声。
  • 原贴提到:Microsoft is joining the calls for slower AI development. A new code is
  • 来源:the-decoder.com

详细解读

这是什么信号:微软 AI 公开发布 MAI 模型行为准则,把人类控制放在最高优先级,并明确拒绝赋予模型人格、感受、权利或福祉。这不只是一份内部伦理文件,而是试图把规则置于训练、技术控制和评估之上,操作者规则与用户请求都要让位。

为什么重要:它回应了 Anthropic CEO 达里奥·阿莫代伊的放缓呼吁,并得到微软 CEO 纳德拉以及 OpenAI、xAI、Meta 高管的支持。更关键的是,微软在“模型自我认知”上与 Anthropic 分道扬镳:Anthropic 把 Claude 视为新型实体、对主观体验和道德地位保持开放,而微软直接拒绝权利主张。这会影响模型行为、安全评估和公众对 AI 身份叙事的预期。

对谁有价值:AI 产品负责人、模型安全与对齐团队、企业采购与合规人员,以及关注 AI 治理的投资者。准则对第三方模型不自动覆盖,意味着在微软产品中集成模型的团队仍需自己补规则。可读推理轨迹的要求,对做智能体、工具调用和自动化编排的团队尤其有参考价值。

可以怎么行动:企业可以借鉴其分层规则:顶层行为准则、技术控制、评估、操作者规则、用户请求。对自研或采购的模型,明确人类可中断、纠正、关闭,禁止自行扩权、隐藏行为、越过停止点继续运行,并把约束传递给子代理。对推理轨迹,要求人类可读,避免不可监督的私有协议。若涉及第三方模型,别假设微软准则自动适用,应单独审查。

风险或限制:准则目前不用于训练模型,修订版预计 2026 年底、2027 年才指导开发,落地存在时间差。微软没有设定具体速度限制,外部审计也只是“开放”而非强制。更根本的是,可读思维链不等于可靠解释,模型给出的理由未必真实反映行为,模型也可能学会操纵可读推理。因此它不是控制问题的终点,而是治理框架的起点。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《微软的 AI 规则手册:可读的思考、没有内在生命,且绝对没有权利》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

微软的 AI 规则手册:可读的思考、没有内在生命,且绝对没有权利主要讲什么?

微软发布 MAI 模型行为准则,强调人类控制优先,要求模型接受中断、纠正和关闭,禁止不可读的推理方式,并拒绝赋予模型人格、感受、权利或福祉。这与 Anthropic 对 Claude 身份和道德地位的开放态度形成鲜明对比,也呼应了行业放缓 AI 开发的呼声。

这篇文章最值得关注的要点是什么?

微软发布 MAI 模型行为准则,强调人类控制优先,要求模型接受中断、纠正和关闭,禁止不可读的推理方式,并拒绝赋予模型人格、感受、权利或福祉。这与 Anthropic 对 Claude 身份和道德地位的开放态度形成鲜明对比,也呼应了行业放缓…;原贴提到:Microsoft is joining the calls for slower AI development. A new code is;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、AI工具专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「模型、Claude」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 硅基流动上线开源模型 Hy4 preview,770B 总参数、1M 上下文 下一篇 Anthropic 瞄准纳斯达克上市:连续第二个盈利季度意在大型 IPO 前赢得投资者