Searcharxiv⌕ Search

arXiv · 2609.29596

How many cards, until the first ace: variations, extensions, lachrymae, confidence, Dirichlets

Abstract

From a deck of cards, how many cards do I need to draw, until the first ace? I identify the distribution for this waiting time $T$, and its satisfyingly nice expected value ${\rm E}\,T=(N+1)/(n+1)$, with $N$ the number of cards and $n$ the number of aces; hence $53/5=10.6$ for the standard setup. After having solved this Question One I go on to certain alternative solutions and extensions, involving e.g. Beta approximations. I also consider the distributions and means for the 2nd, the 3rd, the 4th occurrences of aces, with generalisations, where there is a Dirichlet distribution in wait for us, with further links to order statistics for the uniform. Furthermore, an apparatus is developed for obtaining estimators and full confidence distributions for applications where one knows the number $n$ of aces, but not the deck size $N$; and correspondingly for inference about the unknown population size $N$ when $n$ is known. If you have 1000 people in a room, and need to interview 11 of them until you've found the first left-handed person, how may left-handed are there in the room -- here we need both an estimate and a clear measure of uncertainty.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nils Lid Hjort. 2026-08-27. How many cards, until the first ace: variations, extensions, lachrymae, confidence, Dirichlets. https://arxiv.org/abs/2609.29596

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Variable selection in linear mixed model meta-regression with suspected interaction effects -- How can tree-based methods help?

Detecting interaction effects (IEs) in meta-regression is challenging, especially when few studies are available and many plausible interactions are considered. In many meta-analyses, interpretability is essential, which limits the use of complex machine learning methods. Tree-based approaches offer a potentially useful compromise, but their role in meta-regression with random effects is not yet well understood. This paper examines how traditional linear and tree-based methods can support variable selection for IEs in random effects meta-regression. We compare test-based and information-criterion-based linear selection procedures with meta-CART approaches. These include single trees and tree ensembles, which combine single meta-CARTs into a stability selection ensemble based on bootstrapped data. All methods are evaluated using a real-world meta-analytic dataset and a simulation study. The data-generating process assumes linear IEs, complemented by settings with simple non-linear interactions. Our results show that under strictly linear interactions, linear selection methods perform as expected and achieve superior performance for IE detection. Tree-based methods are more conservative when the number of studies is small, but become competitive as sample size increases, particularly the stability-selected variants. When IEs deviate from strict linearity, Wald-type approaches can deteriorate, whereas criterion-based methods maintain a relatively stable performance. Tree-based methods, particularly stabilized versions of random-effects meta-CART, provide a robust alternative when a sufficient number of observations is available. They could thus be used for pre-selection and sensitivity analyses, guarding against more complex data structures. Additionally, selection frequency patterns from meta-CART tree ensembles can help to reveal structural patterns in the data in an exploratory way.

stat.OT↗

Asymptotic confidence intervals for the difference and the ratio of the weighted kappa coefficients of two diagnostic tests subject to a paired design

The weighted kappa coefficient of a binary diagnostic test is a measure of the beyond-chance agreement between the diagnostic test and the gold standard, and depends on the sensitivity and specificity of the diagnostic test, on the disease prevalence and on the relative importance between the false positives and the false negatives. This article studies the comparison of the weighted kappa coefficients of two binary diagnostic tests subject to a paired design through confidence intervals. Three asymptotic confidence intervals are studied for the difference between the parameters and five other intervals for the ratio. Simulation experiments were carried out to study the coverage probabilities and the average lengths of the intervals, giving some general rules for application. A method is also proposed to calculate the sample size necessary to compare the two weighted kappa coefficients through a confidence interval. A program in R has been written to solve the problem studied and it is available as supplementary material. The results were applied to a real example of the diagnosis of malaria.

stat.OT↗

Engaging students with statistics through choice of real data context on homework

Statistics educators recommend teaching with real data with relevant contexts, but defining relevancy is challenging and varies by student. We investigated whether providing student choice of data context increases engagement through a quasi-experiment in two sections of an introductory probability and statistics course at a large public university (n=65 consenting students). Sections alternated as treatment and control: during their treatment, students chose weekly homework from three similar instructor-provided options varying by data context; during control weeks, they received randomly assigned contexts. We found no significant difference in homework grades between treatment and control conditions. However, thematic analysis revealed students with choice reported enhanced engagement and motivation, greater appreciation for statistics' real-world value, and increased autonomy. Students overwhelmingly preferred contexts relevant to their interests, experiences, daily lives, and career paths-though preferences varied considerably across individuals. Based on these findings, we provide four recommendations for statistics educators: (1) use real data with authentic contexts, (2) select contexts students care about, (3) incorporate variety across data contexts, and (4) consider choice as a pedagogical tool.

stat.OT↗