Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Why safer AI models still harbor hidden gender discrimination
Safety updates to GPT models don't eliminate gender discrimination—they disguise it. Researchers analyzed 450,000 text completions and found that explicit sexual violence against women disappeared in newer models, but was replaced by subtler harms: women's output became 36% less diverse in topic, men gained positive traits women didn't, and harmful content like framing breast cancer as a men's rights issue now scores as non-toxic by standard safety measures.
Companies use toxicity scores to prove their AI models are getting safer with each update. This work shows those scores are misleading. A model can appear cleaner by standard measures while actually distributing gender-based harms more effectively—just in forms that existing safety tools don't catch. This means current methods for auditing AI fairness are insufficient, and organizations need better ways to detect discrimination that evolves rather than disappears.