AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-05-24 5 浏览 公开

趋势解读:Why you shouldn't leave model selection on default,聚焦形式化数学证明能力

实验表明,Microsoft Copilot在默认自动模式下分析文本数据时会编造针对特定国家的刻板印象,而非基于实际数据。该问题源于模型选择不当,推理模型可正确执行任务,但用户需手动切换。这揭示了依赖AI默认设置的风险。

SOURCE / 全球热点解读 MIN / 4 ACCESS / 公开 POST / 2026-05-24 18:17:46

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

An experiment shows that Microsoft Copilot makes up country-specific stereotypes when analyzing text data instead of actually looking at what the data says. In tests using simulated answers about career goals, the AI in standard mode claimed Italians were more interested in art than Brits. The problem: the underlying datasets for both countries were identical. The experiment ran Copilot in "Auto" mode, which is supposed to pick the best model for a given task. It didn't. Reasoning models handled the task just fine, but users need to know how and when to switch to a reasoning model depending on the tool. Most users likely don't. An experiment shows how Microsoft's AI assistant Copilot applies stereotypes when analyzing data instead of actually reading it. Thinking models solve the task but sometimes need users to know their tools. Microsoft Copilot has become the go-to tool for quick data analysis at many companies. But an experiment by mathematician Adam Kucharski shows that when analyzing text data, the tool can spit out results that have nothing to do with the actual data. Instead, it falls back on stereotypes baked into the underlying language model. For the test, Kucharski created 2,000 simulated free-text responses about emotions and labeled them "UK." He then copied the same 2,000 responses and labeled them "US." The combined 4,000 entries were shuffled and handed to Copilot in "Auto" mode for analysis. Ad The result: Copilot delivered a detailed summary of how US and UK respondents supposedly differed. "Based on the dataset you shared, US and UK responses differ mainly in tone, intensity, and wording style, even though they express similar emotional states," the tool concluded. But the data was identical. Ad DEC_D_Incontent-1 In a second experiment, Kucharski pushed harder. He had a language model generate 200 statements about career goals and copied the dataset five times for the US, UK, France, Germany, and Italy. Copilot again produced country-specific differences: Italians were three times more likely to show interest in arts careers than Brits, and Americans were 1.5 times more business-oriented than the French. All five groups contained the same clichéd and biased statements. Ad When Kucharski asked Copilot to dig deeper, the tool first ran a simple keyword-based count. As expected, it returned identical results for all countries. But Copilot ignored its own finding. Instead, it offered a quantified analysis that once again showed made-up differences, this time with completely fabricated percentages. The analysis ran in "Auto" mode, which Microsoft says should pick the best model on its own. It obviously didn't. Most users probably stick with this default in Copilot and in other tools too. The version Kucharski tested is the standard Copilot that comes with a Microsoft 365 Business account. The majority of Copilot users most likely run this version. Ad DEC_D_Incontent-2 "Which means there’s a real risk that people are currently using AI to produce analysis that bears no resemblance to what people actually said," Kucharski writes. If these kinds of analyses were applied to real datasets, groups with no actual differences could end up looking worlds apart, all because of the language model's built-in assumptions about demographic groups. Ad

中文翻译

一项实验表明,Microsoft Copilot 在分析文本数据时会编造针对特定国家/地区的刻板印象,而不是实际查看数据的内容。在使用有关职业目标的模拟答案的测试中,标准模式下的人工智能声称意大利人比英国人对艺术更感兴趣。问题是:两个国家的基础数据集是相同的。该实验在“自动”模式下运行 Copilot,该模式应该为给定任务选择最佳模型。事实并非如此。推理模型可以很好地处理任务,但用户需要知道如何以及何时根据工具切换到推理模型。大多数用户可能不会。一项实验展示了微软的人工智能助手 Copilot 在分析数据而不是实际读取数据时如何应用刻板印象。思维模型可以解决任务,但有时需要用户了解他们的工具。 Microsoft Copilot 已成为许多公司快速数据分析的首选工具。但数学家 Adam Kucharski 的一项实验表明,在分析文本数据时,该工具可能会输出与实际数据无关的结果。相反,它依赖于底层语言模型中的刻板印象。在测试中,库查斯基创建了 2,000 个有关情绪的模拟自由文本响应,并将它们标记为“英国”。然后,他复制了同样的 2,000 条回复,并将其标记为“美国”。合并后的 4,000 个条目被打乱并以“自动”模式交给 Copilot 进行分析。结果:Copilot 详细总结了美国和英国受访者的差异。该工具总结道:“根据您分享的数据集,美国和英国的反应主要在语气、强度和措辞风格上有所不同,尽管它们表达了相似的情绪状态。”但数据是相同的。 在第二个实验中,库查斯基更加努力。他让一个语言模型生成 200 条关于职业目标的陈述,并将美国、英国、法国、德国和意大利的数据集复制了五次。 Copilot 再次产生了针对具体国家的差异:意大利人对艺术职业表现出兴趣的可能性是英国人的三倍,而美国人对商业的兴趣是法国人的 1.5 倍。所有五个团体都包含相同的陈词滥调和偏见言论。当 Kucharski 要求 Copilot 进行更深入的挖掘时,该工具首先运行一个简单的基于关键字的计数。正如预期的那样,它为所有国家/地区返回了相同的结果。但副驾驶忽略了自己的发现。相反,它提供了量化分析,再次显示了虚构的差异,这次是完全捏造的百分比。该分析在“自动”模式下运行,微软表示该模式应该自行选择最佳模型。显然没有。大多数用户可能也会在 Copilot 和其他工具中坚持使用此默认设置。 Kucharski 测试的版本是 Microsoft 365 商业帐户附带的标准 Copilot。大多数 Copilot 用户很可能运行此版本。 “这意味着人们目前使用人工智能进行的分析与人们实际所说的毫无相似之处,这确实存在风险,”库查斯基写道。如果将这些类型的分析应用于真实的数据集,没有实际差异的群体最终可能会显得天壤之别,这都是因为语言模型对人口统计群体的内置假设。

核心信息

实验表明,Microsoft Copilot在默认自动模式下分析文本数据时会编造针对特定国家的刻板印象,而非基于实际数据。该问题源于模型选择不当,推理模型可正确执行任务,但用户需手动切换。这揭示了依赖AI默认设置的风险。

  • Copilot默认模式编造国家刻板印象
  • 自动模式选错模型导致错误结论
  • 推理模型可正确执行但需手动切换
  • 用户使用默认设置存在高风险
  • 需验证AI输出与实际数据一致性

详细解读

这是什么信号: 该实验直接揭示了当前主流AI工具(如Microsoft Copilot)在默认“自动”模式下存在严重的模型选择缺陷。自动模式未能为文本分析任务自动切换到合适的推理模型,反而依赖语言模型中的文化刻板印象输出虚假结论。这不仅是一个技术bug,更是产品设计中对用户使用场景考虑不足的信号。

为什么重要: 许多企业已将Copilot等AI助手集成到数据分析工作流中,默认设置被广泛应用。如果默认模式会产生与数据无关的偏见结论,那么基于这些结论的决策(如市场策略、用户画像)可能完全错误,导致资源浪费和声誉风险。此外,该问题具有普遍性,其他AI工具的默认模型也可能存在类似缺陷。

对谁有价值: 首先,使用AI进行数据分析的企业员工和数据科学家,他们需要意识到不能盲目信任默认设置。其次,AI工具的产品经理和开发者,此实验提示他们在设计模型选择逻辑时需更智能,或给用户提供明确指引。最后,关注AI伦理的群体,因为刻板印象的自动生成会放大社会偏见。

可以怎么行动: 用户在Copilot或其他AI工具中应主动选择“推理模型”(如GPT-4 Turbo with Vision或特定推理模型)来处理文本分析任务,而非依赖“自动”模式。企业可以建立内部AI使用规范,要求对关键数据分析结果进行人工验证,或对AI输出进行交叉检查。开发者在产品中可增加模型选择建议或警告,提示用户当前任务可能不适合默认模型。

风险或限制: 本实验基于模拟数据,现实场景中数据差异可能更复杂,但问题本质相同。风险在于用户可能不知情地继续使用默认模式,且工具厂商可能未及时修复。限制是实验仅测试了Copilot,其他AI工具的默认模式表现可能不同,但原理相似。

信息差价值

信息差价值: 多数AI用户认为默认设置已足够智能,但此实验揭示了一个隐蔽的认知盲区——自动模式可能因模型选择不当而输出完全错误的结论。这一发现挑战了“AI足够可靠”的普遍信任,为行业和用户提供了第一手风险案例。

业务启发: 企业在部署AI工具时,不能仅依赖厂商的默认配置。应识别高频使用场景(如文本数据分析),针对性选择模型或建立规则。同时,此事件也提示AI产品设计者,需在“易用性”与“准确性”之间寻找平衡,例如在自动模式下加入任务分类逻辑,或向用户透明展示当前所用模型及其适用性。

可沉淀动作: ① 制定企业内部AI使用指南,明确不同任务推荐模型;② 在数据分析流程中增加“AI输出与原始数据分布对比”步骤;③ 定期组织员工培训,提升对AI局限性的认知;④ 向工具供应商反馈此类问题,推动产品改进。

参考来源

上一篇 Claude Code自动模式:多任务并行的关键技巧 下一篇 趋势解读:Anthropic may keep supplying Claude to the NSA,提升开发者接入体验