AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
Why AI watermarks fail when lawyers need them most
Three widely-used watermarking systems for AI-generated text collapse under realistic attacks and fail legal evidence standards. When texts were slightly rewritten—a simple paraphrase—nearly every watermarked passage lost its identifying mark entirely: 100% loss for two methods, 98.3% for a third. Even before any attack, false-negative rates were already high (70–83%), meaning the systems missed AI-generated text they should have caught.
Governments are betting on watermarks to prove AI authorship in court and comply with laws like the EU AI Act and California's SB 942. These results show that current watermarks won't hold up under cross-examination or in forensic analysis—a lawyer could easily paraphrase a watermarked text to strip the mark, and a judge applying legal evidence standards would likely reject the watermark as proof at all. This creates a gap between what the law assumes watermarks can do and what they actually can.