✦ For YouGeopoliticsTechFinanceHealthEnergySportsCulture◆ SN Last Week★ Saved

LLM Watermarks Shift Refusals and Tool Calls — SkimNews

By Hacker News · Summarized & edited by · 2026-09-26
LLM Watermarks Shift Refusals and Tool Calls

Get the Tech newsletter

Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.

Why it matters: For developers building agents on watermarked models like Anthropic's future Claude releases, the finding that 16.8% of phi-4 tool verdicts flip between conditions means aggregate accuracy tests can hide consequential wrong-argument or wrong-tool calls. The gemma-3-27b compliance jump under injection — 23.5% churn and a +12.5-point net compliance shift — ties a provenance feature mandated by EU AI Act Article 50(2) directly to a measurable weakening of refusal behavior against adversarial prompts.

Share this story

Ask SkimNews
More tech → Read original →

Get the Tech newsletter

Curated tech stories, every morning. Free.

No spam. Unsubscribe anytime.