觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-09-16 2 浏览 免费阅读

AI 实验室面临数据信任问题,其政策尚未解决

大型企业因数据保留政策收紧而限制使用 Anthropic 旗舰模型 Fable,要求零数据保留保证。英伟达、Booz Allen Hamilton、Palantir 等采取不同应对,凸显 AI 实验室在数据信任与企业 IP 保护上的结构性矛盾。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-09-16 01:52:05

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Major companies are restricting their use of the most advanced AI models or demanding ironclad guarantees that no data gets stored. Nvidia and defense contractor Booz Allen Hamilton are among the companies limiting how they use Anthropic's new flagship model Fable for sensitive work, according to a report from The Information . The pushback started after a policy change in June, when Anthropic said it would retain usage logs from Fable for 30 days to defend against "complex and novel attacks." For companies worried about their intellectual property, that's a dealbreaker. Nvidia now uses Fable only for less sensitive tasks like open-source projects, according to The Information. For internal work like AI-powered supply chain monitoring, the company runs its own Nemotron models instead. "As a company, you know, we believe ZDR [Zero Data Retention] should be on by default," Justin Boitano, Nvidia's VP of Enterprise AI, told The Information. Nvidia has invested in Anthropic, reportedly plans to continue doing so , and supplies the company with hardware for model development. Booz Allen Hamilton, one of the earliest users of Anthropic's Mythos model, has banned employees from using Fable for work on proprietary cybersecurity software, according to The Information. CTO Bill Vass said, "We worry a little bit that [Fable] might be learning from some of our code." Palantir is blocking Fable deployment through its own software to customers until Anthropic grants irrevocable zero-data-retention guarantees, The Information reports. CEO Alex Karp said at a customer event that companies are tired of being "exploited" by AI labs. Karp has been vocal about his distrust before. Palantir's stance is also self-serving, since the company wants customers running AI models through its supposedly secure platform rather than going directly to providers. After customer pushback and OpenAI's August move to let GPT-5.6 Cyber customers store security logs on their own servers , Anthropic followed suit with a similar program rolling out to select customers this fall. Even with zero data retention, labs can learn from how their services get used. Both OpenAI and Anthropic collect metadata and technical usage data from enterprise customers, according to The Information. OpenAI calls this data "de-identified," meaning it's stripped of information that could be traced back to individual customers. OpenAI states on its website that it runs business data through automated classifiers and security tools "to better understand how our services are used." The resulting classifications are metadata about the business data "but do not contain any of the business data itself," the company writes. Some customers aren't sure what exactly that metadata covers, according to The Information, and don't think the current transparency is enough. John Schulman , OpenAI co-founder who briefly worked at Anthropic and now works at Thinking Machines , recently laid out the different ways AI companies can train on user data. The spectrum runs from direct pretraining on user data, which carries a high risk of reproducing content, to distilling large models into smaller ones, to building reinforcement learning tasks from "user traces." That last approach has a low risk of content reproduction but can still extract customer IP. It ranges from harmless ("use explicit user feedback in reward model training") to invasive ("upload user's coding environment and commit history to turn into rl envs"), Schulman says. "De-identification is weak," he adds, and users can be traced back "with just a small number of bits" and it doesn't protect against IP leakage. AI researcher Sarah Hooker, who previously worked at Cohere and Google DeepMind, describes a similar loophole . There are "clever synthetic data techniques that can generate distributional equivalent data while preserving privacy." In other words, even if an AI lab doesn't use original data directly, it might be able to extract

中文翻译

大型公司正在限制其使用最先进 AI 模型的方式,或要求获得绝对保证,确保不存储任何数据。据 The Information 报道,英伟达和国防承包商 Booz Allen Hamilton 属于限制如何使用 Anthropic 新旗舰模型 Fable 处理敏感工作的公司之列。这一抵制始于 6 月的一项政策变更,当时 Anthropic 表示,它将保留 Fable 的使用日志 30 天,以防御“复杂而新型的攻击”。对于担心自身知识产权的公司来说,这是不可接受的。

据 The Information 报道,英伟达现在仅将 Fable 用于开源项目等不那么敏感的任务。对于 AI 驱动的供应链监控等内部工作,该公司转而运行自己的 Nemotron 模型。“作为一家公司,你知道,我们认为 ZDR(零数据保留)应该默认开启,”英伟达企业 AI 副总裁 Justin Boitano 告诉 The Information。英伟达已投资 Anthropic,据报道计划继续投资,并为其模型开发提供硬件。

据 The Information 报道,Booz Allen Hamilton 是 Anthropic 的 Mythos 模型最早的用户之一,它已禁止员工使用 Fable 处理专有网络安全软件的工作。首席技术官 Bill Vass 表示:“我们有点担心 [Fable] 可能正在从我们的某些代码中学习。”据 The Information 报道,Palantir 阻止通过其自有软件向客户部署 Fable,直到 Anthropic 授予不可撤销的零数据保留保证。首席执行官 Alex Karp 在一次客户活动上表示,公司已经厌倦了被 AI 实验室“剥削”。Karp 此前就曾公开表达过他的不信任。Palantir 的立场也符合自身利益,因为该公司希望客户通过其所谓的安全平台运行 AI 模型,而不是直接去找提供商。

在客户抵制以及 OpenAI 在 8 月允许 GPT-5.6 Cyber 客户将安全日志存储在自己的服务器上之后,Anthropic 也采取了类似做法,于今年秋季向部分客户推出了一项类似计划。即使有零数据保留,实验室也可以从其服务的使用方式中学习。据 The Information 报道,OpenAI 和 Anthropic 都从企业客户那里收集元数据和技术使用数据。OpenAI 称这些数据是“去标识化的”,这意味着它被剥离了可追溯到单个客户的信息。OpenAI 在其网站上表示,它通过自动分类器和安全工具运行业务数据,“以更好地了解我们的服务是如何被使用的”。该公司写道,由此产生的分类是关于业务数据的元数据,“但不包含任何业务数据本身”。据 The Information 报道,一些客户不确定这些元数据到底涵盖什么,也不认为当前的透明度足够。

John Schulman,OpenAI 联合创始人,曾短暂在 Anthropic 工作,现在在 Thinking Machines 工作,最近阐述了 AI 公司训练用户数据的不同方式。其范围从直接在用户数据上进行预训练(这带来很高的内容再现风险),到将大型模型蒸馏为较小的模型,再到从“用户痕迹”构建强化学习任务。Schulman 说,最后一种方法的内容再现风险很低,但仍可能提取客户知识产权。它的范围从无害(“在奖励模型训练中使用明确的用户反馈”)到侵入性(“上传用户的编码环境和提交历史,以转化为 rl 环境”)。他补充说,“去标识化很弱”,用户“仅用少量比特”就能被追溯,而且它不能防止知识产权泄露。AI 研究员 Sarah Hooker,此前曾在 Cohere 和 Google DeepMind 工作,描述了一个类似的漏洞。存在“巧妙的合成数据技术,可以在保护隐私的同时生成分布等价的数据”。换句话说,即使 AI 实验室不直接使用原始数据,它也可能能够提取

核心信息

大型企业因数据保留政策收紧而限制使用 Anthropic 旗舰模型 Fable,要求零数据保留保证。英伟达、Booz Allen Hamilton、Palantir 等采取不同应对,凸显 AI 实验室在数据信任与企业 IP 保护上的结构性矛盾。

  • 大型企业因数据保留政策收紧而限制使用 Anthropic 旗舰模型 Fable,要求零数据保留保证。英伟达、Booz Allen Hamilton、Palantir 等采取不同应对,凸显 AI 实验室在数据信任与企业 IP 保护上的结构性矛盾。
  • 原贴提到:Major companies are restricting their use of the most advanced AI models
  • 来源:the-decoder.com

详细解读

这是什么信号:企业客户不再默认接受 AI 实验室的数据保留政策,而是用使用边界和采购决策表达不信任。Anthropic 因 6 月将 Fable 使用日志保留 30 天而遭到英伟达、Booz Allen Hamilton、Palantir 等客户反制,OpenAI 也在 8 月允许 GPT-5.6 Cyber 客户自行存储安全日志,Anthropic 随后跟进。这说明数据治理已从合规附属条款上升为模型商业化的核心变量。

为什么重要:当模型能力差距缩小,数据信任成为企业采购的一级门槛。零数据保留不只是法律条款,而是企业是否愿意把核心代码、供应链、网络安全等敏感工作流交给第三方模型的前提。但文章也指出,即便零数据保留,实验室仍可收集元数据和技术使用数据;OpenAI 称之为“去标识化”,但部分客户不清楚其边界。John Schulman 直言“去标识化很弱”,Sarah Hooker 也提到合成数据可生成分布等价数据,说明技术层面存在间接提取 IP 的空间。

对谁有价值:对企业 AI 采购负责人、CTO、CISO 和安全团队,这是一份采购风险清单;对 AI 平台厂商和投资者,这是理解企业客户流失与信任成本的关键案例;对开发者与超级个体,则提示在选择模型时要把数据条款当作产品能力的一部分。

可以怎么行动:企业应要求不可撤销的零数据保留保证,并明确元数据收集范围;对敏感任务采用私有模型、本地推理或专属部署,对非敏感任务再用前沿模型;在合同中加入日志自主存储、审计权和数据删除条款;建立模型供应商数据条款对比清单,并定期复核。Palantir 的做法虽有其商业动机,但“通过安全平台托管模型”的思路值得参考。

风险或限制:零数据保留可能削弱模型供应商防御复杂攻击的能力,Anthropic 正是以安全防御为由保留日志;完全隔离可能牺牲模型能力与成本效率;供应商仍可能通过元数据、合成数据或强化学习任务间接学习客户 IP。因此,企业不能把信任完全外包给政策承诺,而要用技术架构和合同条款双重兜底。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《AI 实验室面临数据信任问题,其政策尚未解决》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

AI 实验室面临数据信任问题,其政策尚未解决主要讲什么?

大型企业因数据保留政策收紧而限制使用 Anthropic 旗舰模型 Fable,要求零数据保留保证。英伟达、Booz Allen Hamilton、Palantir 等采取不同应对,凸显 AI 实验室在数据信任与企业 IP 保护上的结构性矛盾。

这篇文章最值得关注的要点是什么?

大型企业因数据保留政策收紧而限制使用 Anthropic 旗舰模型 Fable,要求零数据保留保证。英伟达、Booz Allen Hamilton、Palantir 等采取不同应对,凸显 AI 实验室在数据信任与企业 IP 保护上的结构性…;原贴提到:Major companies are restricting their use of the most advanced AI models;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、AI工具专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「模型」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 谷歌发布 Gemini 3.8 Live,以极低成本对标 OpenAI GPT-Live-1 下一篇 Claude for Small Business 新增 43 个工作流和 27 个集成,并推出免费培训计划