Search arXivSearch

arXiv · 2212.09868

Quantifying fairness and discrimination in predictive models

Abstract

The analysis of discrimination has long interested economists and lawyers. In recent years, the literature in computer science and machine learning has become interested in the subject, offering an interesting re-reading of the topic. These questions are the consequences of numerous criticisms of algorithms used to translate texts or to identify people in images. With the arrival of massive data, and the use of increasingly opaque algorithms, it is not surprising to have discriminatory algorithms, because it has become easy to have a proxy of a sensitive variable, by enriching the data indefinitely. According to Kranzberg (1986), "technology is neither good nor bad, nor is it neutral", and therefore, "machine learning won't give you anything like gender neutrality `for free' that you didn't explicitely ask for", as claimed by Kearns et a. (2019). In this article, we will come back to the general context, for predictive models in classification. We will present the main concepts of fairness, called group fairness, based on independence between the sensitive variable and the prediction, possibly conditioned on this or that information. We will finish by going further, by presenting the concepts of individual fairness. Finally, we will see how to correct a potential discrimination, in order to guarantee that a model is more ethical

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Arthur Charpentier. 2022-12-19. Quantifying fairness and discrimination in predictive models. https://arxiv.org/abs/2212.09868

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Changes-in-Changes for Ordered Choice Models with Underreporting

We develop a Difference-in-Differences framework for discrete, ordered outcomes subject to underreporting. Such outcomes commonly arise in self-reported surveys on socially undesirable or stigmatized behaviors, where respondents may conceal their true behavior. For a discrete Changes-in-Changes model that is shown to admit an equivalent threshold-crossing representation, we derive nonparametric bounds for the counterfactual and factual outcome distributions as well as for the associated quantile treatment effects when outcomes are underreported. These bounds are shown to be sharp uniformly across outcome levels under additional support conditions, and we propose suitable estimation and bootstrap inference procedures. In an extension, we also consider a semiparametric underreporting model that allows to point identify and estimate distributional treatment effects. As an application, we investigate the impact of recreational marijuana legalization on the consumption behavior of 8th-grade students in several U.S. states.

econ.EM

Triple Difference Designs with Heterogeneous Treatment Effects

Triple difference designs have become increasingly popular in empirical economics. The advantage of a triple difference design is that, within a treatment group, it allows another subgroup of the population -- potentially less impacted by the treatment -- to serve as a comparison for the subgroup of interest. While literature on difference-in-differences has discussed heterogeneity in treatment effects between treated and control groups or over time, relatively little attention has been given to triple difference designs and the implications of heterogeneity in treatment effects in this setting. In this paper, I show that the parameter identified under common triple difference assumptions does not allow for causal interpretation of differences between subgroups when subgroups may differ in their underlying (unobserved) treatment effects. I propose a new parameter of interest, the controlled difference in average treatment effects on the treated, which allows for causal comparisons between subgroups. I then propose identification assumptions and doubly-robust estimators for this parameter. I use a simulation study to highlight the desirable finite-sample properties of these estimators, as well as to show the difference between the two parameters. An empirical application shows the importance of considering treatment effect heterogeneity in practical applications.

econ.EM

Identification in Linear Quantile Panel Models

This paper studies identification in linear quantile panel models with unrestricted individual heterogeneity when the number of time periods is fixed and small. We impose strict exogeneity, whereby the conditional quantile restriction holds given the individual's complete regressor history and latent individual effect, but otherwise allow the disturbances to be arbitrarily dependent over time.

econ.EM