Augmented Hypothesis Testing with Persona-Based LLM Simulations
Using AI predictions to run cheaper, faster experiments without sacrificing accuracy
Researchers developed a new statistical framework that uses predictions from AI models—including those from persona-based language models—to shrink the sample sizes needed for A/B tests while maintaining statistical rigor. The method works even when the AI predictions are imperfect, automatically downweighting unreliable signals while extracting useful information from accurate ones. In experiments on real datasets, the approach reduced the cost of testing by substantial margins without compromising the validity of results.
A/B testing is the standard way companies validate decisions, but it's slow and expensive—you need to wait weeks or months and recruit thousands of users. This framework could let companies run valid experiments with 50% fewer participants or shorter timelines by feeding in cheap AI predictions alongside real user data. The method is robust to bad predictions, so teams can experiment with this approach even if they're uncertain how good their models are.