CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation
Which AI safety filters actually work best when combined together
Summary
Researchers tested 19 different jailbreak attacks against 15 AI safety defenses to find the best way to stack them. No single defense blocked all attacks, but combining the right ones across different stages of the AI system blocked most attacks without making the AI less useful.
Why it matters
As AI systems become more widely deployed, companies need concrete guidance on how to defend them against adversaries trying to trick them into harmful outputs. This work shows how to layer defenses in order of effectiveness, helping engineers make practical decisions about which safety tools to deploy where in their AI pipelines.