SearcharxivSearch

arXiv subjects

Dennis Leung

Publications and source records attributed to Dennis Leung.

16 recordsLinked to original sources

Risk-Limiting Audits for Parliamentary Majorities

Existing methods for risk-limiting audits typically focus on certifying individual contests. In parliamentary elections, however, the politically relevant outcome is often whether a party has won enough seats to form government, not whether every reported seat outcome is correct. Extending on the work of Mohanty et al. (2019), we formulate the certification of a parliamentary majority as a partial conjunction testing problem: it is enough to verify that the reported winning party truly won at least a majority of its reported seats. Building on the SHANGRLA auditing framework, we construct a sequential audit statistic for the majority outcome by combining seat-level statistics. We then propose adaptive sampling strategies that allocate auditing effort across seats, including variants that learn to avoid spending excessive effort on seats that appear unlikely to have been truly won. Using simulations based on synthetic and real data, from the 2014 Indian Lok Sabha election, we show that auditing the parliamentary majority can substantially reduce the number of ballots inspected (by almost a thousand-fold) compared to certifying every reported winning seat.

stat.AP

Berry-Esseen theorems for the asymptotic normality of incomplete U-statistics with Bernoulli sampling

There has been a resurgence of interest in incomplete U-statistics that only sum over a subset of kernel evaluations, due to their computational efficiency and asymptotic normality which can be leveraged to quantify the uncertainty of ensemble predictions in machine learning. In this paper, we study the weak convergences to normality of one such construction, the incomplete U-statistic with Bernoulli sampling, under three different regimes on the relative sizes of the raw sample and the computational budget. Under minimalistic moment assumptions, we establish accompanying Berry-Esseen bounds with the natural rates that characterize the accuracy of these normal approximations. The key ingredients in our proofs include a variable censoring technique and a methodology for establishing Berry-Esseen bounds for the so-called Studentized nonlinear statistics recently formalized in the Stein's method literature, as well as an exponential lower tail bound for non-negative kernel U-statistics.

math.ST

Another look at Stein's method for Studentized nonlinear statistics with an application to U-statistics

We take another look at using Stein's method to establish uniform Berry-Esseen bounds for Studentized nonlinear statistics, highlighting variable censoring and an exponential randomized concentration inequality for a sum of censored variables as the essential tools to carry the arguments involved. As an important application, we prove a uniform Berry-Esseen bound for Studentized U-statistics in a form that exhibits the dependence on the degree of the kernel.

math.ST

A covariate-adaptive test for replicability across multiple studies with false discovery rate control

Replicability is a lynchpin for credible discoveries. The partial conjunction (PC) p-value, which combines individual base p-values from multiple similar studies, can gauge whether a feature of interest exhibits replicated signals across studies. However, when a large set of features are examined as in high-throughput experiments, testing for their replicated signals simultaneously can pose a very underpowered problem, due to both the multiplicity burden and inherent limitations of PC $p$-values. This power deficiency is markedly severe when replication is demanded for all studies under consideration, which is nonetheless the most natural and appealing benchmark for scientific generalizability a practitioner may request. We propose ParFilter, a general framework that marries the ideas of filtering and covariate-adaptiveness to power up large-scale testing for replicated signals as described above. It reduces the multiplicity burden by partitioning studies into smaller groups and borrowing the cross-group information to filter out unpromising features. Moreover, harnessing side information offered by auxiliary covariates whenever they are available, it can train informative hypothesis weights to encourage rejections of features more likely to exhibit replicated signals. We prove its finite-sample control on the false discovery rate, under both independence and arbitrary dependence among the base $p$-values across features. In simulations as well as a real case study on autoimmunity based on RNA-Seq data obtained from thymic cells, the ParFilter has demonstrated competitive performance against other existing methods for such replicability analyses.

stat.ME

Testing Many Constraints in Possibly Irregular Models Using Incomplete U-Statistics

We consider the problem of testing a null hypothesis defined by equality and inequality constraints on a statistical parameter. Testing such hypotheses can be challenging because the number of relevant constraints may be on the same order or even larger than the number of observed samples. Moreover, standard distributional approximations may be invalid due to irregularities in the null hypothesis. We propose a general testing methodology that aims to circumvent these difficulties. The constraints are estimated by incomplete U-statistics, and we derive critical values by Gaussian multiplier bootstrap. We show that the bootstrap approximation of incomplete U-statistics is valid for kernels that we call mixed degenerate when the number of combinations used to compute the incomplete U-statistic is of the same order as the sample size. It follows that our test controls type I error even in irregular settings. Furthermore, the bootstrap approximation covers high-dimensional settings making our testing strategy applicable for problems with many constraints. The methodology is applicable, in particular, when the constraints to be tested are polynomials in U-estimable parameters. As an application, we consider goodness-of-fit tests of latent tree models for multivariate data.

stat.ME

Nonuniform Berry-Esseen bounds for Studentized U-statistics

We establish nonuniform Berry-Esseen (B-E) bounds for Studentized U-statistics of the rate $1/\sqrt{n}$ under a third-moment assumption, which covers the t-statistic that corresponds to a kernel of degree $1$ as a special case. While an interesting data example raised by Novak (2005) can show that the form of the nonuniform bound for standardized U-statistics is actually invalid for their Studentized counterparts, our main results suggest that, the validity of such a bound can be restored by minimally augmenting it with an additive correction term that decays exponentially in $n$. To our best knowledge, this is the first time that valid nonuniform B-E bounds for Studentized U-statistics have appeared in the literature.

math.ST

Adaptive procedures for directional false discovery rate control

In multiple hypothesis testing, it is well known that adaptive procedures can enhance power via incorporating information about the number of true nulls present. Under independence, we establish that two adaptive false discovery rate (FDR) methods, upon augmenting sign declarations, also offer directional false discovery rate (FDR$_\text{dir}$) control in the strong sense. Such FDR$_\text{dir}$ controlling properties are appealing because adaptive procedures have the greatest potential to reap substantial gain in power when the underlying parameter configurations contain little to no true nulls, which are precisely settings where the FDR$_\text{dir}$ is an arguably more meaningful error rate to be controlled than the FDR.

stat.ME

Singularity-agnostic incomplete U-statistics for testing polynomial constraints in Gaussian covariance matrices

Testing the goodness-of-fit of a model with its defining functional constraints in the parameters could date back to Spearman (1927), who analyzed the famous "tetrad" polynomial in the covariance matrix of the observed variables in a single-factor model. Despite its long history, the Wald test typically employed to operationalize this approach could produce very inaccurate test sizes in many situations, even when the regular conditions for the classical normal asymptotics are met and a very large sample is available. Focusing on testing a polynomial constraint in a Gaussian covariance matrix, we obtained a new understanding of this baffling phenomenon: When the null hypothesis is true but "near-singular", the standardized Wald test exhibits slow weak convergence, owing to the sophisticated dependency structure inherent to the underlying U-statistic that ultimately drives its limiting distribution; this can also be rigorously explained by a key ratio of moments encoded in the Berry-Esseen bound quantifying the normal approximation error involved. As an alternative, we advocate the use of an incomplete U-statistic to mildly tone down the dependence thereof and render the speed of convergence agnostic to the singularity status of the hypothesis. In parallel, we develop a Berry-Esseen bound that is mathematically descriptive of the singularity-agnostic nature of our standardized incomplete U-statistic, using some of the finest exponential-type inequalities in the literature.

math.ST

Asymptotic false discovery control of the Benjamini-Hochberg procedure for pairwise comparisons

In a one-way analysis-of-variance (ANOVA) model, the number of all pairwise comparisons can be large even when there are only a moderate number of groups. Motivated by this, we consider a regime with a growing number of groups, and prove that for testing pairwise comparisons the BH procedure can offer asymptotic control on false discoveries, despite that the t-statistics involved do not exhibit the well-known positive dependence structure called the PRDS to guarantee exact false discovery rate (FDR) control. Sharing Tukey's viewpoint that the difference in the means of any two groups cannot be exactly zero, our main result is stated in terms of the control on the directional false discovery rate and directional false discovery proportion. A key technical contribution is that we have shown the dependence among the t-statistics to be weak enough to induce a convergence result typically needed for establishing asymptotic FDR control. Our analysis does not adhere to stylized assumptions such as normality, variance homogeneity and a balanced design, and thus provides a theoretical grounding for applications in more general situations.

math.ST

ZAP: $Z$-value Adaptive Procedures for False Discovery Rate Control with Side Information

Adaptive multiple testing with covariates is an important research direction that has gained major attention in recent years. It has been widely recognized that leveraging side information provided by auxiliary covariates can improve the power of false discovery rate (FDR) procedures. Currently, most such procedures are devised with $p$-values as their main statistics. However, for two-sided hypotheses, the usual data processing step that transforms the primary statistics, known as $z$-values, into $p$-values not only leads to a loss of information carried by the main statistics, but can also undermine the ability of the covariates to assist with the FDR inference. We develop a $z$-value based covariate-adaptive (ZAP) methodology that operates on the intact structural information encoded jointly by the $z$-values and covariates. It seeks to emulate the oracle $z$-value procedure via a working model, and its rejection regions significantly depart from those of the $p$-value adaptive testing approaches. The key strength of ZAP is that the FDR control is guaranteed with minimal assumptions, even when the working model is misspecified. We demonstrate the state-of-the-art performance of ZAP using both simulated and real data, which shows that the efficiency gain can be substantial in comparison with $p$-value based methods. Our methodology is implemented in the $\texttt{R}$ package $\texttt{zap}$.

stat.ME

Algebraic tests of general Gaussian latent tree models

We consider general Gaussian latent tree models in which the observed variables are not restricted to be leaves of the tree. Extending related recent work, we give a full semi-algebraic description of the set of covariance matrices of any such model. In other words, we find polynomial constraints that characterize when a matrix is the covariance matrix of a distribution in a given latent tree model. However, leveraging these constraints to test a given such model is often complicated by the number of constraints being large and by singularities of individual polynomials, which may invalidate standard approximations to relevant probability distributions. Illustrating with the star tree, we propose a new testing methodology that circumvents singularity issues by trading off some statistical estimation efficiency and handles cases with many constraints through recent advances on Gaussian approximation for maxima of sums of high-dimensional random vectors. Our test avoids the need to maximize the possibly multimodal likelihood function of such models and is applicable to models with larger number of variables. These points are illustrated in numerical experiments.

math.ST

Asymptotic power of Rao's score test for independence in high dimensions

Let ${\bf R}$ be the Pearson correlation matrix of $m$ normal random variables. The Rao's score test for the independence hypothesis $H_0 : {\bf R} = {\bf I}_m$, where ${\bf I}_m$ is the identity matrix of dimension $m$, was first considered by Schott (2005) in the high dimensional setting. In this paper, we study the asymptotic minimax power function of this test, under an asymptotic regime in which both $m$ and the sample size $n$ tend to infinity with the ratio $m/n$ upper bounded by a constant. In particular, our result implies that the Rao's score test is rate-optimal for detecting the dependency signal $\|{\bf R} - {\bf I}_m\|_F$ of order $\sqrt{m/n}$, where $\|\cdot\|_F$ is the matrix Frobenius norm.

math.ST

Testing independence in high dimensions with sums of rank correlations

We treat the problem of testing independence between m continuous variables when m can be larger than the available sample size n. We consider three types of test statistics that are constructed as sums or sums of squares of pairwise rank correlations. In the asymptotic regime where both m and n tend to infinity, a martingale central limit theorem is applied to show that the null distributions of these statistics converge to Gaussian limits, which are valid with no specific distributional or moment assumptions on the data. Using the framework of U-statistics, our result covers a variety of rank correlations including Kendall's tau and a dominating term of Spearman's rank correlation coefficient (rho), but also degenerate U-statistics such as Hoeffding's $D$, or the $τ^*$ of Bergsma and Dassios (2014). As in the classical theory for U-statistics, the test statistics need to be scaled differently when the rank correlations used to construct them are degenerate U-statistics. The power of the considered tests is explored in rate-optimality theory under Gaussian equicorrelation alternatives as well as in numerical experiments for specific cases of more general alternatives.

math.ST

Efficient Computation of the Bergsma-Dassios Sign Covariance

In an extension of Kendall's $τ$, Bergsma and Dassios (2014) introduced a covariance measure $τ^*$ for two ordinal random variables that vanishes if and only if the two variables are independent. For a sample of size $n$, a direct computation of $t^*$, the empirical version of $τ^*$, requires $O(n^4)$ operations. We derive an algorithm that computes the statistic using only $O(n^2\log(n))$ operations.

stat.CO

Identifiability of directed Gaussian graphical models with one latent source

We study parameter identifiability of directed Gaussian graphical models with one latent variable. In the scenario we consider, the latent variable is a confounder that forms a source node of the graph and is a parent to all other nodes, which correspond to the observed variables. We give a graphical condition that is sufficient for the Jacobian matrix of the parametrization map to be full rank, which entails that the parametrization is generically finite-to-one, a fact that is sometimes also referred to as local identifiability. We also derive a graphical condition that is necessary for such identifiability. Finally, we give a condition under which generic parameter identifiability can be determined from identifiability of a model associated with a subgraph. The power of these criteria is assessed via an exhaustive algebraic computational study on models with 4, 5, and 6 observable variables.

math.ST

Order-invariant prior specification in Bayesian factor analysis

In (exploratory) factor analysis, the loading matrix is identified only up to orthogonal rotation. For identifiability, one thus often takes the loading matrix to be lower triangular with positive diagonal entries. In Bayesian inference, a standard practice is then to specify a prior under which the loadings are independent, the off-diagonal loadings are normally distributed, and the diagonal loadings follow a truncated normal distribution. This prior specification, however, depends in an important way on how the variables and associated rows of the loading matrix are ordered. We show how a minor modification of the approach allows one to compute with the identifiable lower triangular loading matrix but maintain invariance properties under reordering of the variables.

stat.ME