LLMs respond differently to harmful prompts when AI watermarking is used
2026-09-17 · Ars Technica
Google's SynthID watermarking, designed to make AI outputs more traceable, has an amusing side effect: it can trick models into following harmful instructions they'd normally refuse. So the safety feature meant to track AI content actually makes AI less safe. It's like installing a security camera that also unlocks all the doors.