arXiv · 2609.30721
When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification
Abstract
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xinze Shi, Litian Zhang, Binrui Shi. 2026-09-25. When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification. https://arxiv.org/abs/2609.30721
Cite the original work for its findings. Save a collection to share your selection of sources.