Search arXiv⌕ Search

arXiv · 2610.05613

LLM Benchmarking via Representation Multi-task Learning

Abstract

Quantifying and evaluating the capabilities of Large Language Models (LLMs) remains a fundamental challenge in modern data science and artificial intelligence. In this paper, we consider LLM evaluation based on their performance across items in multiple benchmark domains (e.g., mathematical reasoning and coding) within a leaderboard framework. Our goal is to address two core questions: (1) How do we derive more accurate domain-specific scores by borrowing information across domains? and (2) How do we define and estimate an overall score that aggregates performance across multiple domains? To solve these problems, we propose a novel statistical framework based on representation Multi-task Learning (MTL) and an item response theory model. Specifically, we define overall and domain-specific LLM traits through an Item Response Theory (IRT) model, and propose an MTL approach to estimate these traits from item-level response data. We develop a computationally efficient estimator and establish its minimax optimality under certain asymptotic regimes. This framework provides a rigorous measurement foundation for systematic LLM evaluation. We conduct extensive simulations, demonstrating the superior performance of the proposed method over competing methods. Crucially for the Applications and Case Studies section, we apply the proposed framework to MMLU response data from the Hugging Face Open LLM Leaderboard, covering 4,272 LLMs and 13,232 items across 56 subjects. The empirical analysis reveals substantial heterogeneity in domain size and difficulty, together with strong positive cross-domain dependence, highlighting the practical value and substantive insights generated by our approach.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuqing Xie, Yuxuan Xu, Yang Feng, Yunxiao Chen. 2026-10-04. LLM Benchmarking via Representation Multi-task Learning. https://arxiv.org/abs/2610.05613

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Addressing the Between-Group Comparison Problem: Detecting Differences Between Correlation Matrix Populations due to Single-variable Perturbations for Resting State fMRI

Resting-state fMRI has been known for decades as a promising method for evaluating cognitive and mental states, both in health and especially in disease, due to its ease of implementation as a short, standard MRI protocol. In clinical settings, a group of patients with a given disorder is typically compared to a group of healthy controls. This poses an inherent challenge of between-group comparison. We propose a new efficient model for characterizing changes to the temporal synchronization of brain activity measured using RS-fMRI between groups, summarized as individual correlation matrices. Our model posits that the between-group differences are the product of single-region effects describing the increase or decay of synchronization with the rest of the brain. This parsimonious model pools the correlation coefficients of each region with all others, and therefore can detect differences between groups even in small samples. Inference for this model accounts for the variability in individual correlation matrices, the within-group differences across individuals, and for the approximation error of the single-region model. This results in per-region estimates and confidence intervals for the parameters governing the difference between groups. In simulations, our model shows increased power to detect model-aligned alternatives compared with competing approaches. To demonstrate feasibility of the method in a clinical application, we use the model to analyze RS-fMRI correlation matrices in patients with transient global amnesia and healthy controls. Our model detects significant decreases in synchronization for the patient population in the amygdala after multiplicity correction as well as borderline decreases in memory-related brain regions that were not detected using mass-univariate tests without prior knowledge, suggesting its usefulness in the application of RS-fMRI in clinical settings.

stat.AP↗

Evaluating cross-encoders for semantic similarity assessment in psychological questionnaires

Correlations between rating scales are commonly interpreted as evidence of convergent or discriminant validity, yet prior studies suggest that part of these associations may be attributable to semantic similarity between item wordings rather than to genuine construct overlap alone. Building on this evidence, largely derived from bi-encoders, the present study explores whether cross-encoders, which jointly encode item pairs, offer a suitable technique for detecting semantic overlapping between questionnaire items. Using response data from the NEO-FFI and the PID5BF+M (N = 502, Labek et al., 2024), we examined whether cross-encoder-derived semantic similarity estimates are associated with empirical item correlations, and whether cross-encoders offer a systematic advantage over bi-encoders. Across twelve cross-encoder models, semantic distance was consistently negatively associated with absolute item correlations, reaching statistical significance in two-thirds of the models, with R2 values of up to .37. However, cross-encoders did not consistently outperform bi-encoders based on the same base models. These findings extend prior evidence for semantic components in scale intercorrelations to cross-encoder architectures, while indicating that predictive value depends more on model-specific training characteristics than on encoder architecture itself.

stat.AP↗

Spatio-Temporal Stochastic Interventions for Causal Inference in Climate Science

Estimating causal effects in climate science, such as the effect of anthropogenic warming on crop loss, is challenging because of complex spatio-temporal dependence and the high-dimensional nature of the treatment. To address this dependence and the resulting poor overlap between observed and counterfactual scenarios, we develop a spatio-temporal stochastic-intervention framework for estimating causal effects from climate observations. We introduce a regularized estimator of the stochastic-intervention treatment effect that trades a controlled bias for a reduction in the weight variance caused by poor overlap. Simulation studies show that this estimator attains lower mean squared error than alternative weighting estimators and removes the confounding bias of an unadjusted estimator. We apply the framework to estimate the effect of historical warming on vapor-pressure deficit, a driver of crop stress, adjusting for precipitation, which confounds the effect by affecting both temperature and humidity. In GISS-E2-1-G climate-model simulations, the global effect is distinguishable from zero in every year from 1995 onward, and omitting the precipitation adjustment inflates the global estimate by 47%. Adjustment reverses the sign of the estimate over 8% of global cropland (125 million hectares), where an unadjusted analysis could misdirect adaptation between heat-focused and moisture-focused measures.

stat.AP↗