From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
How AI safety research shifted from explaining models to controlling them
Over six years, the field studying trustworthy AI moved away from trying to understand why existing models make decisions and toward actively steering new generative systems like ChatGPT to be truthful and safe. Truthfulness research jumped from nearly zero papers in 2021 to over one-third by 2026, while older methods for explaining black-box models declined—then resurged through new mechanistic approaches that peer into how models actually work.
As AI systems become more powerful and widely deployed, researchers need practical ways to keep them honest and safe rather than just understanding them after the fact. This shift shows the field is moving faster than the capabilities themselves, which is necessary if AI developers are to stay ahead of potential harms. The fact that all trust dimensions lit up simultaneously when ChatGPT arrived suggests future model releases will face immediate scrutiny on multiple fronts—something the research community is now better equipped to provide.