Anthropic's Claude Watermark Works by Skewing Word Choices

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic confirmed its Claude text watermarking works via semantic steganography — biasing each token decision toward a secret-key-determined 'green list' of words rather than hiding invisible Unicode characters
- Anthropic's original support document claimed the watermark is 'imperceptible' and 'doesn't change the meaning, quality, or readability' of responses — a claim Gruber argues is contradicted by a technique that systematically steers word selection
- The watermark applies to all Claude text longer than 200 tokens (~150 words), including private one-on-one conversations that will never be publicly shared or scrutinized
- Detection confidence scales with text length — short passages cannot be reliably flagged, while longer texts can be identified with greater statistical certainty based on green-list skew
- Only Anthropic can detect Claude's watermarks because each provider holds its own secret key — Gemini's marks are invisible to Claude and vice versa, fragmenting detection across the industry
- The EU's 'Code of Practice on Transparency of AI-Generated Content' also requires providers to mandate in their terms of service that users not remove watermarks — which, taken literally, could forbid users from rephrasing outputs since word choices themselves are the marks
Why it matters: What Anthropic marketed as an invisible provenance tool actually imposes a systematic constraint on every word choice in any Claude response over 200 tokens, including private chats. Users who prize precision in LLM output now face a quality tax imposed by EU regulatory compliance, with only Anthropic holding the secret key capable of reading its own marks. The contradiction between Anthropic's 'imperceptible' claim and the actual token-biasing technique exposes a gap between the company's marketing and its implementation.
Ask SkimNews


