TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables
Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN 2.5 has the highest point estimate, closely followed by TabPFN 3 and Logistic Regression. A paired bootstrap over the target pool separates RealTabPFN 2.5 from TabPFN 3 by 87 Elo (95\% interval [43, 129]). Tabular foundation models generally occupy the leading ranks, while the strongest configuration depends on the operating point and biomedical modality. The AutoML framework AutoGluon, using its one-hour "extreme" preset, is configured as a separate resource-intensive reference and reported here at the reference and full cell. Fold-level predictions, run status, and deterministic aggregations make every reported result reproducible and reusable. The benchmark is open to contributions of new biomedical datasets. The interactive leaderboard is available at https://tabbench-bio.eu.