AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-08-16 0 浏览 会员

Anthropic公布Claude新水印机制的更多细节

Anthropic发布博客解释其Claude文本水印的工作原理、抗编辑性及对代码的影响,以符合欧盟AI法案透明度要求,但用户反应不一。

SOURCE / AI小生意项目库 MIN / 4 ACCESS / 会员 POST / 2026-08-16 02:58:39

原贴

查看原文
作者:Anthony Ha 来源站点:techcrunch.com 原贴时间:

原文

Anthropic published a blog post Friday seeking to answer some basic questions about how it will watermark the text generated by its chatbot Claude. Such as: How will the watermarking actually work? Can it be hidden with editing? And how does this affect code? Claude users have been debating the move since the company revealed earlier this week that it would be doing this watermarking to comply with the EU AI Act’s Transparency Code, which requires AI companies to use systems that make it possible to identify AI-generated content. On Reddit, for example, one poster characterized this as a conspiracy against innocent Claude users , while another claimed, “The only reason you wouldn’t want this is to lie to people.” And Business Insider reports that “dozens” of users on X have claimed to cancel their Claude subscriptions as a result. Anthropic’s new post starts with a general overview of the watermarking concept, explaining that when making “low-stakes choices” — like choosing between the words “overcast” and “grey” to describe the weather — Claude can create a pattern in its responses that is “undetectable to the reader, but is detectable to anyone who has a key that encodes it.” “Watermarking does not impact the quality of Claude’s output,” the company said. “To a reader, a watermarked response is indistinguishable from an unwatermarked one.” More specifically, Anthropic said it will be using the SynthID-Text approach that the Google DeepMind team outlined in 2024 , and that it plans to release a watermark detection API. It also noted that watermarking is distinct from the AI detection approaches offered by companies like Pangram that look for “tells” in the writing (like the construction “his isn’t [X], it’s [Y]”) to reveal AI usage: “Picking up on these patterns is fundamentally different from checking for a watermark.” Could someone just rewrite the text to hide the watermark? Anthropic said it’s possible, but “light editing probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.” “In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated,” the company said. As for whether the watermark will be detectable in text that was only proofread or edited by Claude, Anthropic said that will depend on “the length of the text and how heavily Claude has edited it.” If it’s only been lightly edited, “nearly all the words” will have been written by the human author and “there’s very little (if anything) for the watermark to attach to.” Code, meanwhile, should have less of a watermark than other text, because the model will need to create working code and won’t have the freedom to choose between a variety of equally valid options. “Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code,” Anthropic said. “But by definition, it will have a negligible effect on the actual code produced.” Anthropic also said that Claude won’t be the only AI chatbot to generate watermarked text, as “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.”

中文翻译

Anthropic周五发布了一篇博客文章,试图回答一些关于如何为聊天机器人Claude生成的文本添加水印的基本问题。例如:水印技术实际上如何运作?能否通过编辑隐藏?以及这对代码有何影响?自本周早些时候公司透露将实施水印以遵守欧盟《人工智能法案》的透明度准则以来,Claude用户一直在争论此举,该准则要求人工智能公司使用能够识别AI生成内容的系统。例如,在Reddit上,一位发帖者将此描述为针对无辜Claude用户的阴谋,而另一位则声称:“你不想这样做的唯一原因就是想对人撒谎。”Business Insider报道称,X上“数十名”用户声称因此取消他们的Claude订阅。

Anthropic的新文章首先对水印概念进行总体概述,解释当做出“低风险选择”时——比如选择用“阴天”还是“灰色”来描述天气——Claude可以在其回复中创造一种“读者无法察觉,但任何拥有编码密钥的人都能检测到”的模式。“水印不影响Claude输出的质量,”公司表示。“对读者而言,带水印的回复与未带水印的回复无法区分。”

更具体地说,Anthropic表示将采用Google DeepMind团队在2024年概述的SynthID-Text方法,并计划发布水印检测API。它还指出,水印与Pangram等公司提供的AI检测方法截然不同,后者通过寻找写作中的“迹象”(如“这不是[X],而是[Y]”的结构)来揭示AI的使用:“捕捉这些模式与检查水印根本不同。”

有人能直接重写文本来隐藏水印吗?Anthropic表示这有可能,但“轻度编辑可能无法完全去除水印”,而“完全重写,替换每一个词,则可以去除。”公司称:“在后一种情况下,当然,文本是否还能被称为AI生成是有争议的。”至于水印是否能在仅由Claude校对或编辑的文本中被检测到,Anthropic表示这取决于“文本的长度以及Claude编辑的程度。”如果只是轻度编辑,“几乎所有单词”都由人类作者撰写,那么“水印几乎没有什么(如果有的话)可以附着。”

与此同时,代码应该比其他文本有更少的水印,因为模型需要生成可运行的代码,无法在多种同等有效的选项中自由选择。“话虽如此,在代码中对特定单词或术语有任意选择的领域,水印可以用上,例如代码中的注释,”Anthropic说。“但根据定义,它对实际生成的代码影响微乎其微。”Anthropic还表示,Claude不会是唯一生成带水印文本的AI聊天机器人,因为“其他主要模型开发商已签署同一份实践准则,并将实施自己的水印。”

核心信息

Anthropic发布博客解释其Claude文本水印的工作原理、抗编辑性及对代码的影响,以符合欧盟AI法案透明度要求,但用户反应不一。

  • Anthropic发布博客解释其Claude文本水印的工作原理、抗编辑性及对代码的影响,以符合欧盟AI法案透明度要求,但用户反应不一。
  • 原贴提到:Anthropic published a blog post Friday seeking to answer some basic ques
  • 来源:techcrunch.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 第二届世界人形机器人运动会新增跳远、举重、拔河等项目,2056台机器人参赛 下一篇 AI生成书籍正淹没亚马逊,并拉低人类作者的单书收入