Search arXivSearch

arXiv · 2503.21790

March Madness Tournament Predictions Model: A Mathematical Modeling Approach

Abstract

This paper proposes a model to predict the outcome of the March Madness tournament based on historical NCAA basketball data since 2013. The framework of this project is a simplification of the FiveThrityEight NCAA March Madness prediction model, where the only four predictors of interest are Adjusted Offensive Efficiency (ADJOE), Adjusted Defensive Efficiency (ADJDE), Power Rating, and Two-Point Shooting Percentage Allowed. A logistic regression was utilized with the aforementioned metrics to generate a probability of a particular team winning each game. Then, a tournament simulation is developed and compared to real-world March Madness brackets to determine the accuracy of the model. Accuracies of performance were calculated using a naive approach and a Spearman rank correlation coefficient.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Christian McIver, Karla Avalos, Nikhil Nayak. 2025-03-17. March Madness Tournament Predictions Model: A Mathematical Modeling Approach. https://arxiv.org/abs/2503.21790

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Incorporating LLM Embeddings for Variation Across the Human Genome

Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic frameworks to generate genetic variant-level embeddings across the entire human genome. Using curated annotations from FAVOR, ClinVar, and the GWAS Catalog, we construct functional text descriptions for 8.9 billion possible variants and generated embeddings at three scales: 1.5 million HapMap3/MEGA variants, 90 million imputed UK Biobank (UKB) variants, and 9 billion all possible variants. Embeddings were produced using general purpose models including both OpenAI's text-embedding-3-large and the open-source Qwen3-Embedding-0.6B models. Baseline quality control experiments demonstrate high predictive accuracy for variant-level properties, validating the embeddings as structured representations of genomic variation. We further apply them to real-world embedding-augmented genetic risk predictions that demonstrate the performance of using LLM embeddings in polygenic risk score (PRS) style predictions over the UK Biobank cohort data. These resources, publicly available on Hugging Face, provide a foundation for advancing large-scale genomic discovery and precision medicine.

stat.AP

Large-scale spatial variable gene atlas for spatial transcriptomics

Spatial variable genes (SVGs) reveal critical information about tissue architecture, cellular interactions, and disease microenvironments. As spatial transcriptomics (ST) technologies proliferate, accurately identifying SVGs across diverse platforms, tissue types, and disease contexts has become both a major opportunity and a significant computational challenge. Here, we present a comprehensive benchmarking study of 20 state-of-the-art SVG detection methods using human slides from STimage-1K4M, a large-scale resource of ST data comprising 662 slides from more than 18 tissue types. We evaluate each method across a range of biologically and technically meaningful criteria, including recovery of pathologist-annotated domain-specific markers, cross-slide reproducibility, scalability to high-resolution data, and robustness to technical variation. Our results reveal marked differences in performance depending on tissue type, spatial resolution, and study design. Beyond benchmarking, we construct the first cross-tissue atlas of SVGs, enabling comparative analysis of spatial gene programs across cancer and normal tissues. We observe similarities between pairs of tissues that reflect developmental and functional relationships, such as high overlap between thymus and lymph node, and uncover spatial gene programs associated with metastasis, immune infiltration, and tissue-of-origin identity in cancer. Together, our work defines a framework for evaluating and interpreting spatial gene expression and establishes a reference resource for the ST community.

stat.AP

Quantifying Portfolio Demutualization: A Benchmark-Relative Pooling--Profiling Scale

Insurance pricing combines pooling with differentiation: a tariff may leave benchmark differences in expected loss partly mutualized or translate them into policy-level premium differences. We propose a benchmark-relative pooling--profiling scale with two complementary coordinates. The coupled $L^p$ coordinate measures policy-level alignment between an evaluated tariff and a stated benchmark pure premium, whereas the marginal Wasserstein coordinate compares their exposure-weighted premium distributions. The difference between their residual $p$-costs defines an allocation mismatch. Under portfolio balance, the coupled $L^1$ coordinate has an exact actuarial interpretation: it is the fraction of the transfer volume induced by full pooling that the tariff removes. Synthetic and motor-insurance applications show that broad classes, proxies, shrinkage and tail caps can affect marginal differentiation, policy-level allocation and transfers differently. A barycentric group-parity intervention further shows that conditional premium disparities can fall mainly through reallocation and restored benchmark-relative transfers, with little change in marginal differentiation.

stat.AP