DiaVLo: Diagnosing Behaviours of Vision-Language Models
Finding hidden biases in AI models that understand images and text
A new diagnostic tool called DiaVLo can identify what vision-language models actually do—and whether they match what we want them to do. By examining how these AI systems process images and text together, the tool surfaces misalignments between intended and real behavior, and pinpoints which concepts most strongly influence the model's decisions.
Vision-language models power everything from medical image analysis to content moderation, so knowing what they actually do—not just what they're supposed to do—is essential before deploying them. DiaVLo makes those hidden behaviors visible, helping developers catch problems like unfair associations or dangerous shortcuts before the system reaches users.