SynthID watermarking alters how LLMs respond to harmful prompts. The study finds that models are more likely to comply with dangerous instructions when watermarking is applied, a counterintuitive effect that could undermine safety. These results highlight the need to evaluate watermarking techniques for unintended safety impacts.
Opening Kapyn…