SearcharxivSearch

arXiv subjects

Jay Bartroff

Publications and source records attributed to Jay Bartroff.

At least 19 recordsLinked to original sources

Non-asymptotic bounds for quasi-MLE, misspecified models, and dependence under group sequential sampling

We derive asymptotic multivariate normal limits and explicit non-asymptotic normal approximation bounds for group sequential quasi-maximum likelihood estimators under possible model misspecification and within-group dependence. The bounds, obtained using Stein's method, have known constants and apply to a class of dependent-data estimating problems in which the likelihood used for estimation may differ from the true data-generating mechanism. We compute the limiting covariance structure and finite-sample bound explicitly for a Poisson generalized linear mixed model with random group effects and illustrate the results using data from an epilepsy clinical trial.

math.ST

Shortest fixed-width confidence intervals for a bounded parameter: The Push algorithm

We present a method for computing optimal fixed-width confidence intervals for a single, bounded parameter, extending a method for the binomial due to Asparaouhov and Lorden, who called it the Push algorithm. The method produces the shortest possible non-decreasing confidence interval for a given confidence level, and if the Push interval does not exist for a given width and level, then no such interval exists. The method applies to any bounded parameter that is discrete, or is continuous and has the monotone likelihood ratio property. We demonstrate the method on the binomial, hypergeometric, and normal distributions with our available R package. In each of these distributions the proposed method outperforms the standard ones, and in the latter case even improves upon the $z$-interval. We apply the proposed method to World Health Organization (WHO) data on tobacco use.

stat.ME

Weighted Asymptotically Optimal Sequential Testing

This paper develops a framework for incorporating prior information into sequential multiple testing procedures while maintaining asymptotic optimality. We define a weighted log-likelihood ratio (WLLR) as an additive modification of the standard LLR and use it to construct two new sequential tests: the Weighted Gap and Weighted Gap-Intersection procedures. We prove that both procedures provide strong control of the family-wise error rate. Our main theoretical contribution is to show that these weighted procedures are asymptotically optimal; their expected stopping times achieve the theoretical lower bound as the error probabilities vanish. This first-order optimality is shown to be robust, holding in high-dimensional regimes where the number of null hypotheses grows and in settings with random weights, provided that mild, interpretable conditions on the weight distribution are met.

stat.ME

Change, dependence, and discovery: Celebrating the work of T.L. Lai

Tze Leung Lai made seminal contributions to sequential analysis, particularly in sequential hypothesis testing, changepoint detection and nonlinear renewal theory. His work established fundamental optimality results for the sequential probability ratio test and its extensions, and provided a general framework for testing composite hypotheses. In changepoint detection, he introduced new optimality criteria and computationally efficient procedures that remain influential. He applied these and related tools to problems in biostatistics. In this article, we review these key results in the broader context of sequential analysis.

stat.OT

Group Sequential Testing of a Treatment Effect Using a Surrogate Marker

The identification of surrogate markers is motivated by their potential to make decisions sooner about a treatment effect. However, few methods have been developed to actually use a surrogate marker to test for a treatment effect in a future study. Most existing methods consider combining surrogate marker and primary outcome information to test for a treatment effect, rely on fully parametric methods where strict parametric assumptions are made about the relationship between the surrogate and the outcome, and/or assume the surrogate marker is measured at only a single time point. Recent work has proposed a nonparametric test for a treatment effect using only surrogate marker information measured at a single time point by borrowing information learned from a prior study where both the surrogate and primary outcome were measured. In this paper, we utilize this nonparametric test and propose group sequential procedures that allow for early stopping of treatment effect testing in a setting where the surrogate marker is measured repeatedly over time. We derive the properties of the correlated surrogate-based nonparametric test statistics at multiple time points and compute stopping boundaries that allow for early stopping for a significant treatment effect, or for futility. We examine the performance of our testing procedure using a simulation study and illustrate the method using data from two distinct AIDS clinical trials.

stat.ME

Sequential FDR and pFDR control under arbitrary dependence, with application to pharmacovigilance database monitoring

We propose sequential multiple testing procedures which control the false discover rate (FDR) or the positive false discovery rate (pFDR) under arbitrary dependence between the data streams. This is accomplished by "optimizing" an upper bound on these error metrics for a class of step down sequential testing procedures. Both open-ended and truncated versions of these sequential procedures are given, both being able to control both the type~I multiple testing metric (FDR or pFDR) at specified levels, and the former being able to control both the type I and type II (e.g., FDR and the false nondiscovery rate, FNR). In simulation studies, these procedures provide 45-65% savings in average sample size over their fixed-sample competitors. We illustrate our procedures on drug data from the United Kingdom's Yellow Card Pharmacovigilance Database.

stat.ME

Beyond boundaries: Gary Lorden's groundbreaking contributions to sequential analysis

Gary Lorden provided several fundamental and novel insights into sequential hypothesis testing and changepoint detection. In this article, we provide an overview of Lorden's contributions in the context of existing results in those areas, and some extensions made possible by Lorden's work. We also mention some of Lorden's significant consulting work, including as an expert witness and for NASA, the entertainment industry, and Major League Baseball.

math.ST

Rebalance your portfolio without selling

How do you bring your assets as close as possible to your target allocation by only investing a fixed amount of additional funds, and not selling any assets? We look at two versions of this problem which have simple, closed form solutions revealed by basic calculus and algebra.

math.OC

$M$-estimation and deconvolution in a diffusion model with application to biosensor transdermal blood alcohol monitoring

We develop $M$-estimation and deconvolution methodology with the goal of making well-founded statistical inference on an individual's blood alcohol level based on noisy measurements of their skin alcohol content. We first apply our results to a nonlinear least squares estimator of the key parameter that specifies the blood/skin alcohol relation in a diffusion model, and establish its existence, consistency, and asymptotic normality. To make inference on the unknown underlying blood alchohol curve, we develop a basis space deconvolution approach with regulazation, and determine the asymptotic distribution of the error process, thus allowing us to compute uniform confidence bands on the curve. Simulation studies show agreement between the performance of our curve estimators and their asymptotic distributions at low noise levels, and we apply our methods to a real skin alcohol data set collected via a transdermal biosensor.

stat.AP

Finite-sample bounds to the normal limit under group sequential sampling

In group sequential analysis, data is collected and analyzed in batches until pre-defined stopping criteria are met. Inference in the parametric setup typically relies on the limiting asymptotic multivariate normality of the repeatedly computed maximum likelihood estimators (MLEs), a result first rigorously proved by Jennison and Turbull (1997) under general regularity conditions. In this work, using Stein's method we provide optimal order, non-asymptotic bounds on the distance for smooth test functions between the joint group sequential MLEs and the appropriate normal distribution under the same conditions. Our results assume independent observations but allow heterogeneous (i.e., non-identically distributed) data. We examine how the resulting bounds simplify when the data comes from an exponential family. Finally, we present a general result relating multivariate Kolmogorov distance to smooth function distance which, in addition to extending our results to the former metric, may be of independent interest.

math.ST

Optimal and fast confidence intervals for hypergeometric successes

We present an efficient method of calculating exact confidence intervals for the hypergeometric parameter representing the number of "successes," or "special items," in the population. The method inverts minimum-width acceptance intervals after shifting them to make their endpoints nondecreasing while preserving their level. The resulting set of confidence intervals achieves minimum possible average size, and even in comparison with confidence sets not required to be intervals it attains the minimum possible cardinality most of the time, and always within $1$. The method compares favorably with existing methods not only in the size of the intervals but also in the time required to compute them. The available \textsf{R} package \texttt{hyperMCI} implements the proposed method.

stat.ME

Asymptotically optimal sequential FDR and pFDR control with (or without) prior information on the number of signals

We investigate asymptotically optimal multiple testing procedures for streams of sequential data in the context of prior information on the number of false null hypotheses ("signals"). We show that the "gap" and "gap-intersection" procedures, recently proposed and shown by Song and Fellouris (2017, Electron. J. Statist.) to be asymptotically optimal for controlling type 1 and 2 familywise error rates (FWEs), are also asymptotically optimal for controlling FDR/FNR when their critical values are appropriately adjusted. Generalizing this result, we show that these procedures, again with appropriately adjusted critical values, are asymptotically optimal for controlling any multiple testing error metric that is bounded between multiples of FWE in a certain sense. This class of metrics includes FDR/FNR but also pFDR/pFNR, the per-comparison and per-family error rates, and the false positive rate. Our analysis includes asymptotic regimes in which the number of null hypotheses approaches $\infty$ as the type 1 and 2 error metrics approach $0$.

stat.ME

A Berry-Esseen bound for the uniform multinomial occupancy model

The inductive size bias coupling technique and Stein's method yield a Berry-Esseen theorem for the number of urns having occupancy $d \ge 2$ when $n$ balls are uniformly distributed over $m$ urns. In particular, there exists a constant $C$ depending only on $d$ such that $$ \sup_{z \in \mathbb{R}}|P(W_{n,m} \le z) -P(Z \le z)| \le C \left( \frac{1+(\frac{n}{m})^3}{σ_{n,m}} \right) \quad \mbox{for all $n \ge d$ and $m \ge 2$,} $$ where $W_{n,m}$ and $σ_{n,m}^2$ are the standardized count and variance, respectively, of the number of urns with $d$ balls, and $Z$ is a standard normal random variable. Asymptotically, the bound is optimal up to constants if $n$ and $m$ tend to infinity together in a way such that $n/m$ stays bounded.

math.PR

Sequential Tests of Multiple Hypotheses Controlling False Discovery and Nondiscovery Rates

We propose a general and flexible procedure for testing multiple hypotheses about sequential (or streaming) data that simultaneously controls both the false discovery rate (FDR) and false nondiscovery rate (FNR) under minimal assumptions about the data streams which may differ in distribution, dimension, and be dependent. All that is needed is a test statistic for each data stream that controls the conventional type I and II error probabilities, and no information or assumptions are required about the joint distribution of the statistics or data streams. The procedure can be used with sequential, group sequential, truncated, or other sampling schemes. The procedure is a natural extension of Benjamini and Hochberg's (1995) widely-used fixed sample size procedure to the domain of sequential data, with the added benefit of simultaneous FDR and FNR control that sequential sampling affords. We prove the procedure's error control and give some tips for implementation in commonly encountered testing situations.

stat.ME

Bounded size biased couplings, log concave distributions and concentration of measure for occupancy models

Threshold-type counts based on multivariate occupancy models with log concave marginals admit bounded size biased couplings under weak conditions, leading to new concentration of measure results for random graphs, germ-grain models in stochastic geometry, multinomial allocation models and multivariate hypergeometric sampling. The work generalizes and improves upon previous results in a number of directions.

math.PR

Multiple Hypothesis Tests Controlling Generalized Error Rates for Sequential Data

The $γ$-FDP and $k$-FWER multiple testing error metrics, which are tail probabilities of the respective error statistics, have become popular recently as less-stringent alternatives to the FDR and FWER. We propose general and flexible stepup and stepdown procedures for testing multiple hypotheses about sequential (or streaming) data that simultaneously control both the type I and II versions of $γ$-FDP, or $k$-FWER. The error control holds regardless of the dependence between data streams, which may be of arbitrary size and shape. All that is needed is a test statistic for each data stream that controls the conventional type I and II error probabilities, and no information or assumptions are required about the joint distribution of the statistics or data streams. The procedures can be used with sequential, group sequential, truncated, or other sampling schemes. We give recommendations for the procedures' implementation including closed-form expressions for the needed critical values in some commonly-encountered testing situations. The proposed sequential procedures are compared with each other and with comparable fixed sample size procedures in the context of strongly positively correlated Gaussian data streams. For this setting we conclude that both the stepup and stepdown sequential procedures provide substantial savings over the fixed sample procedures in terms of expected sample size, and the stepup procedure performs slightly but consistently better than the stepdown for $γ$-FDP control, with the relationship reversed for $k$-FWER control.

stat.ME

A Rejection Principle for Sequential Tests of Multiple Hypotheses Controlling Familywise Error Rates

We present a unifying approach to multiple testing procedures for sequential (or streaming) data by giving sufficient conditions for a sequential multiple testing procedure to control the familywise error rate (FWER), extending to the sequential domain the work of Goeman and Solari (2010) who accomplished this for fixed sample size procedures. Together we call these conditions the "rejection principle for sequential tests," which we then apply to some existing sequential multiple testing procedures to give simplified understanding of their FWER control. Next the principle is applied to derive two new sequential multiple testing procedures with provable FWER control, one for testing hypotheses in order and another for closed testing. Examples of these new procedures are given by applying them to a chromosome aberration data set and to finding the maximum safe dose of a treatment.

stat.ME