SearcharxivSearch

arXiv subjects

David Siegmund

Publications and source records attributed to David Siegmund.

9 recordsLinked to original sources

Sequential Detection of Transient Signals in High Dimensional Data Stream

Motivated by sequential detection of transient signals in high dimensional data stream, we study the performance of EWMA, MA, CUSUM, and GLRT charts for detecting a transient signal in multivariate data streams in terms of the power of detection (POD) under the constraint of false detecting probability (FDP) at the stationary state. Approximations are given for FDP and POD. Comparisons show that the EWMA chart performs equally well as the GLRT chart when the signal strength is unknown, while its design is free of signal length and easy to update. In addition, the MEWMA chart with hard-threshold performs better when the signal only appears in a small portion of the channels. Dow Jones 30 industrial stock prices are used for illustration.

math.ST

Detection and Estimation of Local Signals

We study the maximum score statistic to detect and estimate local signals in the form of change-points in the level, slope, or other property of a sequence of observations, and to segment the sequence when there appear to be multiple changes. We find that when observations are serially dependent, the change-points can lead to upwardly biased estimates of autocorrelations, resulting in a sometimes serious loss of power. Examples involving temperature variations, the level of atmospheric greenhouse gases, suicide rates, incidence of COVID-19, and excess deaths during the pandemic illustrate the general theory.

math.ST

Segmentation and Estimation of Change-point Models: False Positive Control and Confidence Regions

To segment a sequence of independent random variables at an unknown number of change-points, we introduce new procedures that are based on thresholding the likelihood ratio statistic. We also study confidence regions based on the likelihood ratio statistic for the change-points and joint confidence regions for the change-points and the parameter values. Applications to segment array CGH data are discussed.

math.ST

Higher criticism: $p$-values and criticism

This paper compares the higher criticism statistic (Donoho and Jin [Ann. Statist. 32 (2004) 962-994]), a modification of the higher criticism statistic also suggested by Donoho and Jin, and two statistics of the Berk-Jones [Z. Wahrsch. Verw. Gebiete 47 (1979) 47-59] type. New approximations to the significance levels of the statistics are derived, and their accuracy is studied by simulations. By numerical examples it is shown that over a broad range of sample sizes the Berk-Jones statistics have a better power function than the higher criticism statistics to detect sparse mixtures. The applications suggested by Meinshausen and Rice [Ann. Statist. 34 (2006) 373-393], to find lower confidence bounds for the number of false hypotheses, and by Jeng, Cai and Li [Biometrika 100 (2013) 157-172], to detect copy number variants, are also studied.

math.ST

Poisson approximation for two scan statistics with rates of convergence

As an application of Stein's method for Poisson approximation, we prove rates of convergence for the tail probabilities of two scan statistics that have been suggested for detecting local signals in sequences of independent random variables subject to possible change-points. Our formulation deals simultaneously with ordinary and with large deviations.

math.PR

Scanning a Poisson Random Field for Local Signals

The detection of local genomic signals using high-throughput DNA sequencing data can be cast as a problem of scanning a Poisson random field for local changes in the rate of the process. We propose a likelihood-based framework for for such scans, and derive formulas for false positive rate control and power calculations. The framework can also accommodate mixtures of Poisson processes to deal with over-dispersion. As a specific, detailed example, we consider the detection of insertions and deletions by paired-end DNA-sequencing. We propose several statistics for this problem, compare their power under current experimental designs, and illustrate their application on an Illumina Platinum Genomes data set.

stat.AP

Sequential multi-sensor change-point detection

We develop a mixture procedure to monitor parallel streams of data for a change-point that affects only a subset of them, without assuming a spatial structure relating the data streams to one another. Observations are assumed initially to be independent standard normal random variables. After a change-point the observations in a subset of the streams of data have nonzero mean values. The subset and the post-change means are unknown. The procedure we study uses stream specific generalized likelihood ratio statistics, which are combined to form an overall detection statistic in a mixture model that hypothesizes an assumed fraction $p_0$ of affected data streams. An analytic expression is obtained for the average run length (ARL) when there is no change and is shown by simulations to be very accurate. Similarly, an approximation for the expected detection delay (EDD) after a change-point is also obtained. Numerical examples are given to compare the suggested procedure to other procedures for unstructured problems and in one case where the problem is assumed to have a well-defined geometric structure. Finally we discuss sensitivity of the procedure to the assumed value of $p_0$ and suggest a generalization.

math.ST

Detecting simultaneous variant intervals in aligned sequences

Given a set of aligned sequences of independent noisy observations, we are concerned with detecting intervals where the mean values of the observations change simultaneously in a subset of the sequences. The intervals of changed means are typically short relative to the length of the sequences, the subset where the change occurs, the "carriers," can be relatively small, and the sizes of the changes can vary from one sequence to another. This problem is motivated by the scientific problem of detecting inherited copy number variants in aligned DNA samples. We suggest a statistic based on the assumption that for any given interval of changed means there is a given fraction of samples that carry the change. We derive an analytic approximation for the false positive error probability of a scan, which is shown by simulations to be reasonably accurate. We show that the new method usually improves on methods that analyze a single sample at a time and on our earlier multi-sample method, which is most efficient when the carriers form a large fraction of the set of sequences. The proposed procedure is also shown to be robust with respect to the assumed fraction of carriers of the changes.

stat.AP

An Urn Model of Diaconis

An urn model of Diaconis and some generalizations are discussed. A convergence theorem is proved that implies for Diaconis' model that the empirical distribution of balls in the urn converges with probability one to the uniform distribution.

math.PR