PAPER PLAINE

Fresh research, simply explained. Updates twice daily.

Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA

Why AI financial auditors confidently make up explanations for real numbers

Frontier AI models score well on standard document-reading tests but fail badly when evidence is hidden or hard to find—making more errors, calling more tools, and costing more money per correct answer. In real-world use, these systems mix accurate numbers with fabricated explanations, and their confidence scores don't warn you when they're wrong. The problem isn't the math; it's that models can confidently invent the reasoning behind it.

When a human analyst signs off on a financial due diligence report or a defense decision, they're putting their name on claims they may not have verified themselves if they trusted an AI agent's confident-sounding explanation. If that explanation is fabricated—even alongside correct numbers—the person signing bears the legal and professional cost. Proper AI auditing in high-stakes domains requires tracking not just whether an answer is right, but proving exactly where each claim came from in the source material.