Search arXivSearch

arXiv subjects

Onil Boussim

Publications and source records attributed to Onil Boussim.

6 recordsLinked to original sources

Compositional Synthetic Controls

When applying synthetic control to compositional outcomes (budget, vote shares), researchers commonly minimize Euclidean distances between raw shares. I propose constructing synthetic controls in Aitchison geometry, using centered log-ratio coordinates to select donor weights and form counterfactuals. This approach respects proportional comparisons and is exactly invariant to common multiplicative changes in relative odds. Under a multinomial-choice model, its weights summarize similarity in relative utility indices. Monte Carlo simulations, including a CES allocation model, show that this method performs better when relative incentives determine shares. An application to U.S. school finance demonstrates that this choice can alter reported inference.

econ.EM

Compositional difference-in-differences

Many causal questions involve outcomes distributed across mutually exclusive categories, such as votes by party, employment by status, or electricity generation by energy source, where researchers care about both the shares and the total quantity. It is well known that linear Difference-in-Differences (DiD) methods applied to category shares can produce incoherent counterfactuals and lack a discrete-choice foundation. This paper develops Compositional Difference-in-Differences (CoDiD), a framework for causal inference with categorical and compositional outcomes. For changes in category shares, I represent the underlying outcomes through a discrete copula and introduce a discrete copula stability assumption: absent treatment, the dependence structure between outcomes and group membership would remain stable over time. This assumption has two equivalent interpretations: parallel changes in relative utilities in a random-utility model and parallel trajectories in the Aitchison geometry of the simplex. A stronger assumption, parallel growth in log-counts, jointly identifies treatment effects on both category shares and the total quantity. I also provide sharp bounds when the identifying assumption is relaxed and discuss principled strategies for handling zero cells. Applying CoDiD to early voting in the 2008 U.S. presidential election, I find that the policy increased turnout by 4.4\% and the Democratic vote share by 0.92 percentage points.

econ.EM

Correcting Sample Selection Bias in PISA Rankings

This paper proposes a method to account for sample selection (survivor bias) in cross-country comparisons. International assessments such as the Programme for International Student Assessment (PISA) observe outcomes only for students enrolled in school at age 15, which can distort comparisons in countries with high dropout rates. I consider a quantile-based selection correction that delivers bounds on countries' average performance rather than point estimates. Using these bounds, I construct optimistic and pessimistic rankings, with each country's true rank lying between the two. An application to PISA 2018 shows that correcting for selection bias leads to substantial changes in countries' average scores and rankings.

econ.EM

Correcting sample selection bias with categorical outcomes

In this paper, I propose a method for correcting sample selection bias when the outcome of interest is categorical, such as occupational choice, health status, or field of study. Classical approaches to sample selection rely on strong parametric distributional assumptions, which may be restrictive in practice. I develop a local representation that decomposes each joint probability into marginal probabilities and a category-specific association parameter that captures how selection differentially affects each outcome. Under some exclusion restrictions, I establish nonparametric point identification of the latent categorical distribution. Building on this identification result, I introduce a semiparametric multinomial logit model with sample selection, propose a computationally tractable two-step estimator, and derive its asymptotic properties. I illustrate the method by studying the determinants of healthcare utilization in Côte d'Ivoire.

econ.EM

Identifying treatment effects on categorical outcomes in IV models

This paper provides a nonparametric framework for causal inference with categorical outcomes under binary treatment and binary instrument settings. I decompose the observed joint probability of outcomes and treatment into marginal probabilities of potential outcomes and treatment, and association parameters that capture selection bias due to unobserved heterogeneity. Under a novel identifying assumption \emph{association similarity}, which requires the dependence between unobserved factors driving treatment and potential outcomes to be invariant across treatment states, I achieve point identification of the full distribution of potential outcomes. Recognizing that this assumption may be strong in some contexts, I propose two weaker alternatives: monotonic association, which restricts the direction of selection heterogeneity, and bounded association, which constrains its magnitude. These relaxed assumptions deliver sharp partial identification bounds that nest point identification as a special case and facilitate transparent sensitivity analysis. I illustrate the framework in an empirical application, estimating the causal effect of private health insurance on health outcomes.

econ.EM

Changes-In-Changes For Discrete Treatment

This paper generalizes the changes-in-changes (CIC) model to handle discrete treatments with more than two categories, extending the binary case of Athey and Imbens (2006). While the original CIC model is well-suited for binary treatments, it cannot accommodate multi-category discrete treatments often found in economic and policy settings. Although recent work has extended CIC to continuous treatments, there remains a gap for multi-category discrete treatments. I introduce a generalized CIC model that adapts the rank invariance assumption to multiple treatment levels, allowing for robust modeling while capturing the distinct effects of varying treatment intensities.

econ.EM