arXiv · 2609.09264
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Abstract
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.
Explore related subjects
Keep this discovery
Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary. 2026-09-08. StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean. https://arxiv.org/abs/2609.09264
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.