Search arXivSearch

arXiv · 2608.14696

The Calibration-Leverage Tradeoff in Exactly Solvable Win-Probability Models

Abstract

We study ball-by-ball win probability (WP) for second-innings run chases in Twenty20 cricket, built as an exactly solvable Markov model: we estimate a single object, the per-ball outcome distribution over {0,...,6, wicket}, and derive WP for every game state by backward induction over the acyclic (balls, wickets, runs-required) chase graph. This construction makes WP an exact martingale, which in turn makes leverage (how much a ball can swing WP) and win probability added (WPA) well-defined and exactly attributable; we use them to confirm that finishers and death bowlers occupy the highest-leverage moments. We then show the model's WP is systematically miscalibrated, and that this is not incidental. Localizing the error, we rule out tail-thinning and marginal mis-estimation (the model's per-ball outcome distribution matches the empirical one to a total variation of at most 0.02 at every required run rate). The only remaining cause is unmodelled dependence given the state, and we identify it: a permutation-null decomposition shows short-range sequential run-scoring persistence (roughly 3-5 balls; innings-level heterogeneity contributes only about 18%; wickets, if anything, anti-cluster). A block-bootstrap simulator that injects the real dependence while holding the marginals fixed closes 26% of the calibration gap (replicated over six seeds), saturating at block lengths of about 20 balls, consistent with the measured correlation range; this is a constructive lower bound that confirms the diagnosis. The result is a structural tradeoff relative to the (balls, wickets, runs) state description: exact leverage requires the martingale, the martingale requires conditional ball independence on that state, and that independence is what miscalibrates the WP. Exactly attributable leverage and well-calibrated WP cannot be obtained from the same object over this state.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Devansh Mishra. 2026-08-09. The Calibration-Leverage Tradeoff in Exactly Solvable Win-Probability Models. https://arxiv.org/abs/2608.14696

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analogy to von Neumann's complexity threshold for self-reproducing automata, we argue that sustainable recursive self-improvement in Large Language Models (LLMs) requires a functional analogue: introspection -- the system's capacity to simulate its own operations and target modifications. Grounded in Kleene's Second Recursion Theorem, we demonstrate the theoretical existence of such introspective programs. However, an empirical review reveals that while current LLMs exhibit quasi-introspection (e.g., partial metacognition), they fall short of true introspection due to structural bottlenecks: a lack of complete self-access, the feedforward nature of the Transformer, and computational class constraints that prevent fixed-point iteration. We conclude by outlining architectural paths to cross this complexity threshold and discussing the associated safety implications.

physics.soc-ph

Multilayer Analysis of the Global Trade Network

Global trade is more than a single network of aggregate flows. Beneath the observable exchange of products among economies lies a complex multilayer structure, formed by thousands of product-specific trade relationships that differ in their similarity, interdependence, and temporal evolution. Using the CEPII's BACI database, which records bilateral product-level trade flows between economies, we represent the global trade network from 1995 to 2024 as a temporal multilayer network, with economies as nodes and directed weighted trade flows as edges. To investigate product-level organisation and cross-layer similarity, temporal structural change, and the structural role of individual economies, we introduce a random-walk-based similarity measure that provides a unified framework for comparing weighted and directed trade layers. Our results show that the global trade network remains relatively stable over short periods but undergoes gradual structural change over longer timescales. We also find that similarity-based product communities only partially align with the official product taxonomy, indicating that products assigned to the same official category do not necessarily exhibit similar trade-network structures. Finally, we show that an economy's structural influence is not always determined by its trade volume. These results highlight the value of multilayer network analysis for revealing patterns in global trade that remain hidden at the aggregate level.

physics.soc-ph

Detectability limits of scaling laws

Power law scaling relations between size and output are central to quantitative theories of cities, organisms, and other complex systems. Competing theories predict scaling exponents that differ by small fractions, but there is no existing theory for verifying whether a given dataset can even distinguish exponents at the required resolution to address such discrepancies. Here we derive a resolution limit for scaling exponents, giving the smallest exponent difference that any method of analysis can detect. We find that the Hurst exponents governing the evolution of systems' sizes and deviations from the scaling law determine how long a record of growing systems must be before it can separate competing scaling theories. Empirical results suggest that many available data panels are insufficient for reliable scaling model selection.

physics.soc-ph