arXiv · 2609.24228
IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts
Abstract
Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural prompts: when stereotype-relevant attributes are left unspecified, models still default to stereotypical outputs. We introduce IMPLICIT-Bench, a benchmark for measuring implicit bias in T2I models under such prompts. The key design is a structured-knowledge-graph (KG) construction of controlled prompt triplets: neutral, stereotype, and anti-stereotype variants that differ only along a single bias dimension while preserving scene semantics. This enables precise attribution of bias effects that template benchmarks cannot achieve. IMPLICIT-Bench comprises 5,493 prompts across 11 bias categories, validated through multi-model agreement, CLIP-based verification, and human evaluation. Using this benchmark, we show that state-of-the-art T2I models exhibit systematic bias under neutral prompts, a failure mode largely invisible to existing evaluations. We then use IMPLICIT-Bench to evaluate debiasing methods, uncovering a fundamental trade-off between bias reduction and semantic fidelity.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yue Dai, Ziyang Liu, Marc Cheong, Caren Han. 2026-09-21. IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts. https://arxiv.org/abs/2609.24228
Cite the original work for its findings. Save a collection to share your selection of sources.