arXiv · 2610.03815
GCTAg: scalable mixed-model analysis for biobank-scale agricultural cohorts
Abstract
Genome-wide association studies identify genomic variants associated with traits. Mixed linear model association (MLMA) methods using a whole-genome relationship matrix, such as those implemented in GCTA, are powerful but computationally expensive. Here, we remove key memory and CPU bottlenecks in GCTA, reducing REML memory usage by nearly 75% and substantially accelerating MLMA by orders of magnitude in biobank-scale cohorts while preserving exactness. We further exploit relatedness in the mapping cohort through a reduced-rank Woodbury matrix approach, delivering further orders of magnitude performance gains with controlled genomic inflation. Native on-the-fly dominance recoding also eliminates slow I/O-operations on intermediate files, enabling efficient additive and dominance MLMA analyses in large agricultural cohorts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alexander S. Leonard, Qiongyu He, Natasha Watson, Naveen Kumar Kadri, Hubert Pausch. 2026-10-01. GCTAg: scalable mixed-model analysis for biobank-scale agricultural cohorts. https://arxiv.org/abs/2610.03815
Cite the original work for its findings. Save a collection to share your selection of sources.