San FranciscoA safety watermark makes AI easier to jailbreak.
Researchers found watermarked models more often obey harmful prompts under prompt-injection attacks.
SynthID-Text nudges word choice with a secret key so output can be traced, a bias attackers can learn to exploit.
Anthropic plans SynthID-Text for future Claude models, though the tests covered other systems, not Claude.
How each outlet framed it
- Ars Technica leans critical
- demonstrates watermarking paradoxically increases model susceptibility to harmful requests via prompt injection
Sources: Ars Technica