Search arXivSearch

arXiv · 2309.07664

Computer says 'no': Exploring systemic bias in ChatGPT using an audit approach

Abstract

Large language models offer significant potential for increasing labour productivity, such as streamlining personnel selection, but raise concerns about perpetuating systemic biases embedded into their pre-training data. This study explores the potential ethnic and gender bias of ChatGPT, a chatbot producing human-like responses to language tasks, in assessing job applicants. Using the correspondence audit approach from the social sciences, I simulated a CV screening task with 34,560 vacancy-CV combinations where the chatbot had to rate fictitious applicant profiles. Comparing ChatGPT's ratings of Arab, Asian, Black American, Central African, Dutch, Eastern European, Hispanic, Turkish, and White American male and female applicants, I show that ethnic and gender identity influence the chatbot's evaluations. Ethnic discrimination is more pronounced than gender discrimination and mainly occurs in jobs with favourable labour conditions or requiring greater language proficiency. In contrast, gender discrimination emerges in gender-atypical roles. These findings suggest that ChatGPT's discriminatory output reflects a statistical mechanism echoing societal stereotypes. Policymakers and developers should address systemic bias in language model-driven applications to ensure equitable treatment across demographic groups. Practitioners should practice caution, given the adverse impact these tools can (re)produce, especially in selection decisions involving humans.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Louis Lippens. 2024-02-12. Computer says 'no': Exploring systemic bias in ChatGPT using an audit approach. https://doi.org/10.1016/j.chbah.2024.100054

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment

Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.

econ.GN

Bricks or Cash? Externalities of Housing Upgrading in High-density Cities

We estimate housing externalities in a high-density city, exploiting the staggered rollout of Singapore's nationwide Main Upgrading Programme for public housing. Controlling for nonrandom neighborhood exposure, we find that upgrading raises treated buildings' prices by 11.5% upon completion and neighboring buildings' resale prices by about 2% within 500 meters, decaying to zero beyond. A model with distance-decaying externalities shows that in dense settings spillovers justify the distortions of in-kind provision; this advantage diminishes and reverses at lower densities. Administrative data on over 2 million residents show that upgrading disproportionately retains older incumbents, suggesting age-specific amenities as an underexplored externality channel.

econ.GN

The Joneses Visit an Economics Lab

Existing literature offers persuasive evidence that individuals care about how their consumption compares to that of peers, and proposes a large variety of explanatory models. The present paper proposes a common framework for many of those models, and compares their ability to predict behavior in a laboratory experiment. We find evidence of Keeping up with the Joneses motivations but also find that conspicuous consumption is enhanced by Veblen motivations arising from peers' ability to observe one's own choice. Among the seven quasi-linear preference models we compare, our data are best explained by a model that contrasts envy and pride (upward vs downward comparisons) using a value function borrowed from Prospect Theory.

econ.GN