FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Can AI tell if two people already know each other from watching them talk?
Researchers created FriendBench, a test that asks whether two people in a 20-second conversation are strangers or already familiar with each other. The best AI systems matched human accuracy across video, audio, and text—but they got there differently: humans weighed both possibilities equally, while AI models were biased toward guessing "stranger." Interestingly, only humans actually benefited from watching body language and facial expressions; AI didn't gain much from video over speech alone.
Detecting familiarity from behavior is crucial for AI assistants that need to navigate social contexts—whether moderating online interactions, analyzing team dynamics, or providing appropriate responses in social settings. The finding that current AI systems misread social cues in systematic ways shows where these models still lag behind humans, and highlights which behavioral channels (like visible nonverbal cues) remain underexploited in multimodal AI training.