QuoteBench: How Matched Scores Can Hide Command-Path Failures
Why AI coding agents' test scores don't measure what actually happens
When AI systems issue computer commands, the score measuring their success can hide massive failures that happen after the AI generates its answer. Researchers built QuoteBench to separate problems caused by the AI itself from problems introduced by how the system processes and executes those commands—and found that one AI model's apparent 3.6-point deficit actually concealed a 64.3-point gap masked by compensating errors elsewhere.
AI coding agents are being deployed to write and run real commands on servers. If their test scores don't reflect actual execution failures, teams deploying these systems won't know when they're truly unreliable. This research shows that standard benchmarks can rank models backwards depending on how commands are processed, potentially leading organizations to trust agents that fail more often than measured.