arXiv · 2609.40124
Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions
Abstract
Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chahat Raj, Sina Mansouri, Aylin Caliskan, Antonios Anastasopoulos, Ziwei Zhu. 2026-09-30. Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions. https://arxiv.org/abs/2609.40124
Cite the original work for its findings. Save a collection to share your selection of sources.