arXiv · 2610.05212
Regression models for ordinal compositional data
Abstract
We present a unified, transformation-free linear regression framework tailored for ordinal compositional data, such as distributions across educational levels or aggregate Likert-scale survey responses. Traditional log-ratio approaches often obscure interpretability, ignore the inherent ordering of categories, and struggle with boundary zero values. To overcome these limitations, we propose a deterministic model constrained to the space of column-stochastic transformation matrices, naturally preserving the fundamental geometry of the simplex. By adopting the weighted 1-Wasserstein distance as the loss function, our method directly embeds the ordinal nature of the data into the estimation process. We provide a computationally efficient and globally optimal solution explicitly formulated as a Linear Programming (LP) problem. The framework systematically addresses four distinct regression scenarios, introducing a novel ordinal tensor product to rigorously handle interactions across multiple compositional predictors. Furthermore, we equip the model with a suitable regularization strategy and novel diagnostic tools, including a Wasserstein-based coefficient of determination ($R^2_W$) and an Order Preservation Index (OPI). Extensive simulation studies and two empirical applications demonstrate the robust predictive performance and interpretability of the proposed methodology.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Beniamino Cappelletti-Montano; Monica Musio; Nicola Piras. 2026-10-04. Regression models for ordinal compositional data. https://arxiv.org/abs/2610.05212
Cite the original work for its findings. Save a collection to share your selection of sources.