PAPER PLAINE

Fresh research, simply explained. Updates twice daily.

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark

Why a popular AI safety test might be measuring multiple things at once

Researchers examined whether a widely-used AI safety benchmark actually measures what it claims to measure: a model's ability to refuse harmful requests. They found that the benchmark's overall score hides the fact that different AI models are being tested on fundamentally different things, making it impossible to fairly compare them. The problem gets worse when results are combined into a single number that obscures these hidden differences.

AI safety benchmarks guide which models get deployed in the real world. If the benchmark score doesn't actually tell us whether a model will safely refuse dangerous requests—because it's secretly measuring multiple unrelated behaviors—then companies and regulators could confidently use models that fail at the specific safety property we thought we were checking. This research flags a systematic problem: collapsing diverse test results into one number can create an illusion of understanding when the underlying measurements are broken.