PAPER PLAINE

Fresh research, simply explained. Updates twice daily.

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Which AI safety filters actually work best when combined together

Researchers tested 19 different jailbreak attacks against 15 AI safety defenses to find the best way to stack them. No single defense blocked all attacks, but combining the right ones across different stages of the AI system blocked most attacks without making the AI less useful.

As AI systems become more widely deployed, companies need concrete guidance on how to defend them against adversaries trying to trick them into harmful outputs. This work shows how to layer defenses in order of effectiveness, helping engineers make practical decisions about which safety tools to deploy where in their AI pipelines.