KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Measuring whether AI can actually use real cybersecurity tools correctly
Current AI systems can answer cybersecurity questions but struggle to translate those answers into working commands for real tools—even tiny syntax errors cause failures. Researchers created KaliBench, a test set of 8,500 real commands across 1,600 cybersecurity tools, and found that no open-source AI model gets more than 42% of commands exactly right without help. Training models on these exact-command tests improved performance dramatically, bringing small models up to the level of much larger ones.
Cybersecurity analysts increasingly rely on AI to automate repetitive tool work, but a single misplaced flag or misspelled argument can break an investigation or miss a threat. This benchmark forces AI systems to prove they can generate working commands, not just sound knowledgeable—which is essential before deploying them in real security operations where mistakes have real consequences.