Searcharxiv⌕ Search

arXiv subjects

William J. Dwyer

Publications and source records attributed to William J. Dwyer.

5 recordsLinked to original sources

Sparse departures from independence in two-way tables: a heteroscedasticity profile and detection boundary, an adaptive higher-criticism gate, and an assumption-lean exact anchor

Two-way contingency tables are tested for independence throughout applied statistics (genomics, network and text co-occurrence, pharmacovigilance, ecology, survey cross-tabulation), and the routine test reads asymptotic Gaussian tail probabilities off the table cell by cell. On the tables people actually analyze this is badly miscalibrated: across 5,543 real public two-way tables, two-thirds have counts small enough or margins heterogeneous enough that the asymptotic maximum-cell independence scan false-positives at a mean 33% against a 0.05 target, while an exact margin-conditional anchor holds at about 1%. The failure is sharpest when the departure is sparse, the association concentrated in a few cells, where the omnibus chi-square is inefficient and the sharp object is the detection boundary. In three parts we answer where a sparse signal can be seen, which combiner attains that limit, and how to calibrate at small counts. Part I specializes the sparse-detection theory of Ingster (1997), Donoho and Jin (2004), and Chhor, Mukherjee and Sen (2024) to independence: the table is a heteroscedastic Gaussian sequence whose variance profile is the expected-count table, fixed by the margins via CVe^2 = (1 + CVr^2)(1 + CVc^2) - 1, giving an exact signal map and a closed-form separation radius valid under a moderate-deviation growing-count condition. Part II shows higher criticism, with its closed-form Jaeschke-Eicker null, attains that boundary adaptively over unknown sparsity. Part III measures the small-count calibration failure (size 0.3 to 0.6 even under uniform margins, so the driver is the per-cell tail), removes it with exact margin-conditional laws (size 0.003 to 0.009), and gives a routing rule sending large-count tables to the asymptotic gate and small-count tables to the exact anchor. Every claim is reproduced from openly deposited, deterministically seeded code.

stat.ME↗

An Honest Effect Size for Contingency Tables: Why Nothing Can Be Unbiased, Where to Put the Error Instead, and How to Route the Report

Cramer's V is the effect size reported beside almost every chi-square test, read against Cohen's labels, yet a large fraction of such numbers describe sampling noise and the standard bias correction (Bergsma 2013) does not fix it. We assemble three results and claim only the consequence of each. First, the squared effect size phi^2 admits no unbiased estimator at any sample size: under fixed-N multinomial sampling the expectation of any estimator is a polynomial in the cell probabilities, while phi^2 is not. Second, unbiasedness transfers across a rescaling of the effect size only if the rescaling is affine, and among affine choices V^2 = phi^2/k is the one bounded in [0,1] with value 1 at perfect association. Third, we give an interval with conservative, asymptotically valid coverage, obtained by projecting a likelihood-ratio confidence set for the cell probabilities through the effect-size map; its lower endpoint is zero in closed form exactly when the test of independence fails to reject. Because nothing is unbiased, the only question is where the irreducible error is placed. Bergsma's correction puts zero error at the null but several percent under the alternative; a delete-one jackknife on the V^2 scale spreads it thin everywhere (absolute bias at most 0.008 over a 180-design grid). Pooling does not remove bias: across 3,000 simulated meta-analyses, pooling 200 studies drives the naive estimator's chance of landing within 0.01 of the truth to zero, while the jackknife's rises to 0.99. The projected interval covers 0.997-1.000 across the tested designs, while a noncentral inversion undercovers (0.936 at phi^2 = 0.18) at two to three times the width. The point estimate and the interval are different problems with different answers; conflating them is why the literature has neither. We give a routing rule and a browser tool that implements it.

stat.ME↗

Exact Conditional Distributions of Chi-Square-Family Statistics for Two-Way Contingency Tables, by Cell-Separable Dynamic Programming

Every chi-square-family statistic for a two-way contingency table (Pearson's X^2, the power-divergence members, the variance-stabilized T_root) is referred to an approximate null distribution (chi-square, moment-matched chi-square, saddlepoint, or bootstrap) that miscalibrates on sparse or heterogeneous tables, where the exact distribution is a lattice of atoms rather than a continuous curve. The exact reference conditions on both margins (the multivariate Fisher noncentral hypergeometric law) but is usually treated as uncomputable, because the number of tables sharing a margin is astronomical. We show that the exact conditional moments and the exact conditional distribution of any chi-square-family statistic are computable without enumerating tables, because the statistic is additive over cells and the margin-conditional law factorizes cell by cell: a dynamic program walks the table one cell at a time, carrying per row-capacity state either a few moment accumulators, a map from statistic value to probability, or one complex number (the characteristic function), at a cost set by the number of margin states rather than the number of tables. The moment engine computes the exact conditional moments of a 5x5 table at three per cell (about 79 billion tables on the margin fibre) in three seconds; the distribution engine returns the exact tail, verified against enumeration to machine precision, where a three-moment chi-square fit misses it by up to 0.28. The engine furnishes Monte-Carlo-free ground truth against which any approximate reference can be measured, extends to the non-null alternative, and is the computational substrate beneath the exact conditional interval and reporting tools of the companion papers.

stat.ME↗

An Exact Noise Floor for Contingency-Table Effect Sizes: A Per-Table Reporting Gate, and How Often It Would Change Reported Magnitudes

Cramer's V and the other chi-square-family effect sizes are positive under exact independence, so a reported value can be smaller than what the same statistic produces on a table with no association at all. The upward bias is classical (Tschuprow 1925; Bartlett 1937; corrected by Bergsma 2013), and reporting the boundary value, the critical effect size, beside the observed one has been recommended in general form (Perugini et al. 2025); we claim neither. Our contribution is two things a general recommendation does not supply. First, an exact per-table reporting gate: the noise floor V_0.95, the value the effect size reaches by chance 5% of the time under independence, computed from the exact margin-conditional law rather than the asymptotic critical value. The asymptotic floor miscalibrates where reporting is riskiest: on heterogeneous tables its false-alarm rate reaches 0.10 against a 0.05 target, while the exact floor holds nominal. Second, a measurement: across 4,129 real two-way tables from 291 public datasets, 21.8% of labeled effects (cluster-robust 95% CI 17 to 27%) sit on tables with no significant association. We recommend reporting V_0.95, in exact form on sparse or heterogeneous tables, beside every association measure and at the design stage.

stat.ME↗

Exact Conditional Confidence Intervals for Cramér's V: Near-Nominal and Tight Where the Guaranteed Interval Is Wide and the Software Interval Does Not Cover

Cramér's V, the effect size reported beside almost every chi-square test, is almost never accompanied by a confidence interval, because the interval is a hard nuisance-parameter problem: infinitely many tables share one effect-size value. The two intervals an analyst can reach for are unsatisfactory. The projection of a joint confidence region onto the effect size is guaranteed for every table but wide, over-covering at essentially 1.000 on larger tables; the noncentral inversion the software prints, and the bootstrap, are narrow but do not cover (median coverage 0.31 on this study's grid). This paper supplies an interval that is both valid and informative. Conditioning on the observed margins makes the exact conditional distribution of the Pearson statistic under a non-null association computable with no asymptotics and no Monte Carlo; inverting it yields a near-nominal confidence interval for Cramér's V that is a third to a half the width of the projection, the advantage growing with table size. The one price, stated plainly, is a change of estimand to the effect size at the observed margins; the mid-p default undercovers small effects, where a conservative variant or the projection remains the fallback.

stat.ME↗