arXiv · 2609.36970
Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages
Abstract
The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different explanations. This creates a critical vulnerability in AutoML: x-hacking, the selective post-hoc choice of a model based on its explanation rather than predictive merit. No existing AutoML framework exposes this risk. We introduce ARSA ML, an open-source Python framework that quantifies Rashomon set structure and predictive multiplicity within AutoML pipelines. Using ARSA ML, we benchmark AutoGluon and H2O across 28 binary classification datasets, and conduct a post-hoc x-hacking analysis revealing a consistent structural asymmetry: AutoGluon produces larger, diverse sets with stable explanations, while H2O generates compact sets with markedly higher prediction divergence and explanation instability -- making H2O users considerably more exposed to x-hacking. This gap persists across all evaluated metrics and epsilon thresholds, pointing to a fundamental difference in each framework's model-building strategy. ARSA ML is available at https://pypi.org/project/arsa-ml/ .
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Katarzyna Woźnica, Katarzyna Rogalska, Zuzanna Sieńko, Mustafa Cavus. 2026-09-29. Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages. https://arxiv.org/abs/2609.36970
Cite the original work for its findings. Save a collection to share your selection of sources.