arXiv · 2609.22637
Orthogonal Policy Learning with Ordinal Outcomes
Abstract
Policy learning methods based on conditional average treatment effects can obscure subpopulation heterogeneity when applied to ordinal outcomes. We develop a policy learning framework for ordinal outcomes with heterogeneous utilities for individuals who strictly benefit from treatment and those who do not. Since the probability of strict benefit is only partially identified without further conditions, we adopt a minimax strategy which minimizes the worst-case regret over the identification region. To estimate the resulting nonsmooth objective, we combine smooth approximations with Neyman orthogonalization to remove first-order bias from nuisance estimation. We also derive an excess worst-case regret bound over a restricted policy class. The proposed method is validated through extensive simulations and an application to the 2022 Survey of Income and Program Participation (SIPP) data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yue Zhang, Shanshan Luo, Yangbo He. 2026-09-18. Orthogonal Policy Learning with Ordinal Outcomes. https://arxiv.org/abs/2609.22637
Cite the original work for its findings. Save a collection to share your selection of sources.