arXiv · 2609.26962
CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
Abstract
Large imbalanced tabular datasets make repeated gradient-boosted tree training expensive. Existing coreset methods often lose accuracy when most majority examples are removed. We present CRISP (Coreset Reduction via Importance-Stratified Pruning), a linear-time method that allocates a negative-class budget across quantile strata of a proxy-model score. Sample weights account for unequal inclusion probabilities. At 95% negative-class reduction on a production fraud dataset, CRISP trains on approximately 1.70M of 25M rows and retains 99.7% of full-data Average Precision. This is a 93.2% reduction in total training rows. On public CriteoPrivateAds, CRISP has the highest mean Average Precision at each tested rate from 90% to 99.4% majority reduction. Sparkov results are mixed at lower rates, but CRISP has the highest mean at 99.2% and 99.4%. Ablations identify budget allocation and inverse-propensity weighting as the main sources of the production-dataset gain.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani. 2026-09-22. CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning. https://arxiv.org/abs/2609.26962
Cite the original work for its findings. Save a collection to share your selection of sources.