SearcharxivSearch

arXiv subjects

Richard Lockhart

Publications and source records attributed to Richard Lockhart.

8 recordsLinked to original sources

SARS-CoV-2 Transmission in University Classes

We investigate transmission dynamics for SARS-CoV-2 on a real network of classes at Simon Fraser University, a medium-sized school in Western Canada. Outbreaks are simulated over the course of one semester across numerous parameter settings for a realistic compartment model, including asymptomatic and presymptomatic transmission. We investigate the control strategy of moving large classes online while small classes are allowed to meet in person. Regression trees are used to model the effect of disease parameters on simulation outputs; specifically, the total number of infections and the peak number of simultaneous cases. We find that an aggressive class size thresolding strategy is required to mitigate the risk of a large outbreak, and that transmission by symptomatic individuals is a key driver of outbreak size,

stat.AP

Cramer-von Mises tests for Change Points

We study two nonparametric tests of the hypothesis that a sequence of independent observations is identically distributed against the alternative that at a single change point the distribution changes. The tests are based on the Cramer-von Mises two-sample test computed at every possible change point. One test uses the largest such test statistic over all possible change points; the other averages over all possible change points. Large sample theory for the average statistic is shown to provide useful p-values much more quickly than bootstrapping, particularly in long sequences. Power is analyzed for contiguous alternatives. The average statistic is shown to have limiting power larger than its level for such alternative sequences. Evidence is presented that this is not true for the maximal statistic. Asymptotic methods and bootstrapping are used for constructing the test distribution. Performance of the tests is checked with a Monte Carlo power study for various alternative distributions.

math.ST

A Generalized Hosmer-Lemeshow Goodness-of-Fit Test for a Family of Generalized Linear Models

Generalized linear models (GLMs) are used within a vast number of application domains. However, formal goodness of fit (GOF) tests for the overall fit of the model$-$so-called "global" tests$-$seem to be in wide use only for certain classes of GLMs. In this paper we develop and apply a new global goodness-of-fit test, similar to the well-known and commonly used Hosmer-Lemeshow (HL) test, that can be used with a wide variety of GLMs. The test statistic is a variant of the HL test statistic, but we rigorously derive an asymptotically correct sampling distribution of the test statistic using methods of Stute and Zhu (2002). Our new test is relatively straightforward to implement and interpret. We demonstrate the test on a real data set, and compare the performance of our new test with other global GOF tests for GLMs, finding that our test provides competitive or comparable power in various simulation settings. Our test also avoids the use of kernel-based estimators, used in various GOF tests for regression, thereby avoiding the issues of bandwidth selection and the curse of dimensionality. Since the asymptotic sampling distribution is known, a bootstrap procedure for the calculation of a p-value is also not necessary, and we therefore find that performing our test is computationally efficient.

stat.ME

A Bayesian Search for the Higgs Particle

The statistical procedure used in the search for the Higgs boson is investigated in this paper. A Bayesian hierarchical model is proposed that uses the information provided by the theory in the analysis of the data generated by the particle detectors. In addition, we develop a Bayesian decision making procedure that combines the two steps of the current method (discovery and exclusion) into one and can be calibrated to satisfy frequency theory error rate requirements. .

stat.AP

A significance test for the lasso

In the sparse linear regression setting, we consider testing the significance of the predictor variable that enters the current lasso model, in the sequence of models visited along the lasso solution path. We propose a simple test statistic based on lasso fitted values, called the covariance test statistic, and show that when the true model is linear, this statistic has an $\operatorname {Exp}(1)$ asymptotic distribution under the null hypothesis (the null being that all truly active variables are contained in the current lasso model). Our proof of this result for the special case of the first predictor to enter the model (i.e., testing for a single significant predictor variable against the global null) requires only weak assumptions on the predictor matrix $X$. On the other hand, our proof for a general step in the lasso path places further technical assumptions on $X$ and the generative model, but still allows for the important high-dimensional case $p>n$, and does not necessarily require that the current lasso model achieves perfect recovery of the truly active variables. Of course, for testing the significance of an additional variable between two nested linear models, one typically uses the chi-squared test, comparing the drop in residual sum of squares (RSS) to a $χ^2_1$ distribution. But when this additional variable is not fixed, and has been chosen adaptively or greedily, this test is no longer appropriate: adaptivity makes the drop in RSS stochastically much larger than $χ^2_1$ under the null hypothesis. Our analysis explicitly accounts for adaptivity, as it must, since the lasso builds an adaptive sequence of linear models as the tuning parameter $λ$ decreases. In this analysis, shrinkage plays a key role: though additional variables are chosen adaptively, the coefficients of lasso active variables are shrunken due to the $\ell_1$ penalty. Therefore, the test statistic (which is based on lasso fitted values) is in a sense balanced by these two opposing properties - adaptivity and shrinkage - and its null distribution is tractable and asymptotically $\operatorname {Exp}(1)$.

math.ST

Exact Post-Selection Inference for Sequential Regression Procedures

We propose new inference tools for forward stepwise regression, least angle regression, and the lasso. Assuming a Gaussian model for the observation vector y, we first describe a general scheme to perform valid inference after any selection event that can be characterized as y falling into a polyhedral set. This framework allows us to derive conditional (post-selection) hypothesis tests at any step of forward stepwise or least angle regression, or any step along the lasso regularization path, because, as it turns out, selection events for these procedures can be expressed as polyhedral constraints on y. The p-values associated with these tests are exactly uniform under the null distribution, in finite samples, yielding exact type I error control. The tests can also be inverted to produce confidence intervals for appropriate underlying regression parameters. The R package "selectiveInference", freely available on the CRAN repository, implements the new inference tools described in this paper.

stat.ME

Methods to distinguish between polynomial and exponential tails

In this article two methods to distinguish between polynomial and exponential tails are introduced. The methods are mainly based on the properties of the residual coefficient of variation for the exponential and non-exponential distributions. A graphical method, called CV-plot, shows departures from exponentiality in the tails. It is, in fact, the empirical coefficient of variation of the conditional excedance over a threshold. The plot is applied to the daily log-returns of exchange rates of US dollar and Japan yen. New statistics are introduced for testing the exponentiality of tails using multiple thresholds. Some simulation studies present the critical points and compare them with the corresponding asymptotic critical points. Moreover, the powers of new statistics have been compared with the powers of some others statistics for different sample size.

stat.ME