arXiv · 2609.27176
Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure
Abstract
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact. We present LeakScale, an interventional framework for estimating this missing quantity. LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-derivable from the public task, controls access to that information, and estimates the resulting control-adjusted change in executable accuracy. Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improves accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points. These findings separate two empirical questions that are often conflated: whether benchmark contact occurred and how strongly a reported score depends on it. LeakScale makes the latter directly measurable.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Divyansh Singh. 2026-09-23. Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure. https://arxiv.org/abs/2609.27176
Cite the original work for its findings. Save a collection to share your selection of sources.