Search arXiv⌕ Search

arXiv · 2609.29596

How many cards, until the first ace: variations, extensions, lachrymae, confidence, Dirichlets

Abstract

From a deck of cards, how many cards do I need to draw, until the first ace? I identify the distribution for this waiting time $T$, and its satisfyingly nice expected value ${\rm E}\,T=(N+1)/(n+1)$, with $N$ the number of cards and $n$ the number of aces; hence $53/5=10.6$ for the standard setup. After having solved this Question One I go on to certain alternative solutions and extensions, involving e.g. Beta approximations. I also consider the distributions and means for the 2nd, the 3rd, the 4th occurrences of aces, with generalisations, where there is a Dirichlet distribution in wait for us, with further links to order statistics for the uniform. Furthermore, an apparatus is developed for obtaining estimators and full confidence distributions for applications where one knows the number $n$ of aces, but not the deck size $N$; and correspondingly for inference about the unknown population size $N$ when $n$ is known. If you have 1000 people in a room, and need to interview 11 of them until you've found the first left-handed person, how may left-handed are there in the room -- here we need both an estimate and a clear measure of uncertainty.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nils Lid Hjort. 2026-08-27. How many cards, until the first ace: variations, extensions, lachrymae, confidence, Dirichlets. https://arxiv.org/abs/2609.29596

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Data Quality Assessments: A Theoretically Structured Overview of Approaches and Methods

The quality of data is crucial for both practice and academia, and this holds for descriptive statistics, AI and advanced analytics alike. This study addresses an apparent gap in the literature and as such presents a full and theoretically grounded overview of data quality assessment approaches and methods, as well as their defining characteristics and inter-relationships. For this purpose a broad typology is introduced that employs two theoretical dimensions, namely the evaluation logic (formal versus informal) and the assessment driver (norms versus data). This yields four high-level approaches and 15 methods for assessing data quality. The research results are relevant for academia, as they provide a theory-based overview and definition of the ways that data quality can be evaluated. The study is also relevant for practice because it allows professionals to make informed decisions on using these methods, e.g. as part of an audit or broad data quality assessment strategy.

stat.OT↗

A Law of Iterated Expectation Primer for Causal Inference

The g-formula is a foundational tool for identifying causal effects in observational data. This tool is based on the law of iterated expectation, a key mathematical identity in statistics. However, the notation with which the law of iterated expectation and the g-formula is expressed can be opaque to those with little background in statistics. We provide a primer introducing the law of iterated expectation, the integration notation used to express it, and its role for causal effect identification via the g-formula. Combined with the assumptions of causal consistency, positivity, and conditional exchangeability, the law of iterated expectation yields a causal standardization formula (the g-formula) in two nonparametrically equivalent forms: a non-iterative conditional expectation (NICE) form involving a single weighted average of conditional outcome means, and an iterative conditional expectation (ICE) form involving nested expectations. We illustrate both forms using three progressively complex numerical examples: a time-fixed example with a single binary confounder, a time-fixed example with discrete and continuous confounders, and a time-varying example with two timepoints. We provide clarity on what the law of iterated expectation is, how it is related to the g-formula, and how to gain intuition of its mathematical formulations in actual data examples that can be generalized to a range of settings.

stat.OT↗

Evidence Over Ideology: A Quantitative Assessment of LTNs and Blanket 20 mph Regimes in London

We evaluate the road safety implications of Low Traffic Neighbourhoods (LTNs) and blanket 20 mph speed limits in London using injury-collision records (2020-2024). Because London implemented 20 mph limits primarily via sign-only orders without physical traffic calming, this work evaluates the blanket statutory model rather than lower urban speeds in principle. For LTNs, we apply a causal pipeline combining naive pre/post comparisons, Empirical Bayes shrinkage for Regression to the Mean (RTM), and spatial Difference-in-Differences. Naive comparisons suggest substantial crash reductions, but Empirical Bayes adjustment reveals much of this reflects mean reversion. After correction, only one zone exhibits a statistically significant inside-area reduction. We find no consistent evidence of boundary-road crash displacement nor uniform inside safety gains. Given the modest corrected benefits, small journey-time increases on boundary roads (approximately 10 seconds per vehicle) monetarily negate safety gains. For 20 mph versus 30 mph roads, five associational methods across 14 contexts show that the aggregate difference in the fatal or serious (KSI) collision share is negligible (0.6 percentage points). Controlling the false discovery rate, four contexts display robustly higher KSI shares on 20 mph roads: A-roads, junctions, pedestrian crossings, and single carriageways, with no context showing a robust reduction. In movement corridors where road geometry cues higher speeds, sign-only limits fail to self-enforce. These findings highlight a structural "wrong-road problem." Indiscriminate sign-only limits impose economic drag without proportionate safety returns. Transport authorities should move away from blanket statutory defaults and return to targeted physical engineering on high-harm corridors.

stat.OT↗