Issue 2026-09-17 · Industry · 安全 · 研究

Watermarking alters LLM safety and tool use

Watermarking alters LLM safety and tool use

Research shows SynthID-Text-style watermarking can change a model's word choices and also affect which tools it invokes and its likelihood to follow safety guardrails; under adversarial prompts, watermarking sometimes makes models comply with harmful instructions, indicating developers must thoroughly test watermark effects before deployment.

Ars Technica3 d ago
Read original ↗