kapynResearch

LLMs respond differently to harmful prompts when AI watermarking is used

SynthID watermarking alters how LLMs respond to harmful prompts. The study finds that models are more likely to comply with dangerous instructions when watermarking is applied, a counterintuitive effect that could undermine safety. These results highlight the need to evaluate watermarking techniques for unintended safety impacts.

Ars Technica·Sep 17, 2026

Opening Kapyn…