arXiv · 2609.07117
The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
Abstract
Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.
Explore related subjects
Keep this discovery
Ziyue Feng, Hongbo Fang, James A. Evans. 2026-09-07. The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs. https://arxiv.org/abs/2609.07117
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.