LLM Watermarks Shift Refusals and Tool Calls — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic announced future Claude models will embed invisible watermarks built on Google DeepMind's SynthID-Text, making behavioral side effects of watermarking a deployed concern for Claude Platform API and cloud-provider customers.
- Article 50(2) of the EU AI Act requires providers of synthetic-text AI systems to mark outputs as machine-readable and artificially generated, giving watermarking regulatory weight beyond voluntary provenance.
- The authors coin "sampling drift" — non-distortionary watermarks still shift token selection under a fixed key, so the same model can answer differently with and without a watermark even though quality is unchanged on average.
- On the Berkeley Function Calling Leaderboard v4, watermarking reduced tool-call accuracy on 6 of 7 tested models (significant on 4), with paired verdict churn averaging 6.5% across 21 model-temperature combinations and bootstrap intervals excluding zero in every case.
- At T=1.0, phi-4 showed 16.8% of call verdicts flipping between watermarked and unwatermarked runs while net accuracy dropped only 2.87 points, while Llama-3.1-8B showed 9.9% churn against a 0.87-point net loss — showing aggregate scores can mask real behavioral shifts.
- Under a fixed prompt-injection technique, gemma-3-27b's refusal churn jumped from 6.0% to 23.5% at T=0.001 and net compliance shifted from −1.0 to +12.5 points, indicating watermarks can make models more compliant under adversarial prompts.
- Models that already over-refuse, such as phi-4 and Qwen3-4B, showed little movement under either condition; the authors warn their limited drift should not be read as evidence that watermarking preserves safety behavior.
Why it matters: For developers building agents on watermarked models like Anthropic's future Claude releases, the finding that 16.8% of phi-4 tool verdicts flip between conditions means aggregate accuracy tests can hide consequential wrong-argument or wrong-tool calls. The gemma-3-27b compliance jump under injection — 23.5% churn and a +12.5-point net compliance shift — ties a provenance feature mandated by EU AI Act Article 50(2) directly to a measurable weakening of refusal behavior against adversarial prompts.
Ask SkimNews




