Auditing Cross-Lingual Fairness in Language Model Watermarking
Why watermarks on AI text work differently across world languages
Watermarking schemes that hide invisible marks in AI-generated text work unevenly across languages, with gaps driven by fundamental linguistic differences rather than random variation. Testing six watermarking methods across eleven languages revealed that performance gaps cluster by language family—not individual languages—suggesting the problem is baked into how these schemes handle different grammatical structures and writing systems.
As AI systems deploy globally, watermarks are a key tool for detecting machine-generated content in moderation and authenticity verification. If watermarks fail silently for speakers of certain languages, some users get strong detection while others face unreliable protection. The framework here provides a concrete way to audit these gaps before deployment, rather than discovering them after problems emerge at scale.