PAPER PLAINE

Fresh research, simply explained. Updates twice daily.

ContractScrub: A benchmark for final review of legal contracts

Testing whether AI can catch hidden mistakes in legal contracts

Researchers created the first test to measure how well AI language models can spot errors in legal contracts—a task lawyers spend hours doing by hand. The results were sobering: even the most advanced models caught fewer than 75% of mistakes, revealing a significant gap between how well these systems perform on general tests and how well they work on real legal documents.

Contract review is expensive, tedious work that consumes thousands of lawyer hours annually. If AI could reliably automate it, firms could cut costs and speed up deals. This benchmark shows that current AI systems aren't ready for the task despite their strong general capabilities—meaning companies relying on these tools for legal review could miss costly errors, and the field needs better, domain-specific AI development before automation is safe to deploy.