Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task
How AI language models struggle when rules conflict with their training instincts
Large language models show the same conflict effects that human brains do when default behaviors clash with explicit rules—six out of seven tested models performed worse when instructions contradicted their ingrained tendencies. Using attention analysis, researchers found that the models deploy different neural pathways depending on whether the rule agrees or conflicts with their defaults, suggesting these effects come from competition between learned patterns and real-time instructions rather than from a unified decision process.
These findings reveal how AI systems make decisions when faced with competing demands, which matters for predicting when they'll follow explicit instructions reliably versus reverting to patterns baked into their training. Understanding this conflict mechanism could improve how we design prompts and fine-tune models to follow rules even when they contradict the model's default behavior—important for safety and reliability in high-stakes applications.