StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Teaching AI to stop dangerous tool use before it happens
Researchers created StepGuard, a safety system that catches risky actions by AI agents right before they execute—like blocking a file deletion or unauthorized data access. The system cuts successful attacks by 77% while barely slowing down the AI's useful work (dropping performance by less than 3%).
AI agents that interact with real systems—reading files, sending emails, modifying databases—pose serious security risks if they go rogue or get hijacked. Current safety checks only look back after damage is done. StepGuard shifts protection to the moment of decision, making it harder for attackers or malfunctioning systems to cause harm without sacrificing the AI's ability to do legitimate work.