SearcharxivSearch

arXiv subjects

Robert Lund

Publications and source records attributed to Robert Lund.

18 recordsLinked to original sources

Genetic Algorithms in Regression

Many statistical problems involve optimization over a discrete parameter space having an unknown dimension. In such settings, gradient-based methods often fail due to the non-differentiability of the objective function or a non-convex or massive search space with an objective function having many local maxima/minima. This paper presents GAReg, a unified genetic algorithm package that handles discrete optimization regression problems, which works well when standard algorithms are unjustified. GAReg provides a compact chromosome representation supporting optimal knot placement for regression splines, best-subset regression variable selection, and related problems. The package allows for uniform initialization, constraint-preserving crossover and mutation, steady-state replacement, and an optional island-model parallelization. GAReg efficiently searches high-dimensional model spaces, providing near-optimal solutions in settings where exhaustive enumeration or integer or dynamic programming approaches are infeasible.

stat.AP

Single Changepoint Procedures

Single changepoint tests have become a staple check for homogeneity of a climate time series, suggesting how climate has changed should non-homogeneity be declared. This paper summarizes the most prominent single changepoint tests used in today's climate literature, relating them to one and other and unifying their presentations. Asymptotic quantiles for the individual tests are presented. Derivations of the quantiles are given, enabling the reader to tackle cases not considered within. Our work here studies both mean and trend shifts, covering the most common settings arising in climatology. SOI and global temperature series are analyzed within to illustrate the techniques.

stat.ME

The Asymptotic Distribution for a Single Joinpoint Changepoint Model

A single joinpoint changepoint model partitions a time series into two segments, joined at the changepoint time by constraining the estimated piecewise linear regression responses to be continuous. This manuscript derives the exact asymptotic distribution of the changepoint existence test statistic gauging whether or not a second segment is necessary. The identified asymptotic distribution, a supremum of a Gaussian process over the unit interval, is rather unwieldy. The work presented here provides the result and its derivation; quantiles of the asymptotic distribution are presented for the user. This addresses a subtle gap in the changepoint literature.

stat.ME

Decoder-only Clustering in Attributed Graphs

This manuscript studies nodal clustering in graphs having multivariate attributes at each node. The framework includes node-specific priors for low-dimensional representations, coupled with a neural decoder that bridges observed attributes with latent variables. Structural and attribute information are incorporated through a graph-fused LASSO regularization on the prior means, promoting nodal clustering. The optimization problem is solved via alternating direction method of multipliers, with Langevin dynamics for posterior inference. Simulation studies on grid graphs, and applications to real data with complex settings, demonstrate the effectiveness of the proposed clustering method.

stat.ME

Is a Recent Surge in Global Warming Detectable?

The global mean surface temperature is widely studied to monitor climate change. A current debate centers around whether there has been a recent (post-1970s) surge/acceleration in the warming rate. This paper addresses whether an acceleration in the warming rate is detectable from a statistical perspective. We use changepoint models, which are statistical techniques specifically designed for identifying structural changes in time series. Four global mean surface temperature records over 1850-2023 are scrutinized within. Our results show limited evidence for a warming surge; in most surface temperature time series, no change in the warming rate beyond the 1970s is detected. As such, we estimate minimum changes in the warming trend for a surge to be detectable in the near future.

stat.AP

Poisson Count Time Series

This paper reviews and compares popular methods, some old and some very recent, that produce time series having Poisson marginal distributions. The paper begins by narrating ways where time series with Poisson marginal distributions can be produced. Modeling nonstationary series with covariates motivates consideration of methods where the Poisson parameter depends on time. Here, estimation methods are developed for some of the more flexible methods. The results are used in the analysis of 1) a count sequence of tropical cyclones occurring in the North Atlantic Basin since 1970, and 2) the number of no-hitter games pitched in major league baseball since 1893. Tests for whether the Poisson marginal distribution is appropriate are included.

stat.ME

High-dimensional latent Gaussian count time series: Concentration results for autocovariances and applications

This work considers stationary vector count time series models defined via deterministic functions of a latent stationary vector Gaussian series. The construction is very general and ensures a pre-specified marginal distribution for the counts in each dimension, depending on unknown parameters that can be marginally estimated. The vector Gaussian series injects flexibility into the model's temporal and cross-dimensional dependencies, perhaps through a parametric model akin to a vector autoregression. We show that the latent Gaussian model can be estimated by relating the covariances of the counts and the latent Gaussian series. In a possibly high-dimensional setting, concentration bounds are established for the differences between the estimated and true latent Gaussian autocovariances, in terms of those for the observed count series and the estimated marginal parameters. The results are applied to the case where the latent Gaussian series is a vector autoregression, and its parameters are estimated sparsely through a LASSO-type procedure.

math.ST

Trends in Northern Hemispheric Snow Presence

This paper develops a mathematical model and statistical methods to quantify trends in presence/absence observations of snow cover (not depths) and applies these in an analysis of Northern Hemispheric observations extracted from satellite flyovers during 1967-2021. A two-state Markov chain model with periodic dynamics is introduced to analyze changes in the data in a grid by grid fashion. Trends, converted to the number of weeks of snow cover lost/gained per century, are estimated for each study grid. Uncertainty margins for these trends are developed from the model and used to assess the significance of the trend estimates. Grids with questionable data quality are identified. Among trustworthy grids, snow presence is seen to be declining in almost twice as many grids as it is advancing. While Arctic and southern latitude snow presence is found to be rapidly receding, other locations, such as Eastern Canada, are experiencing advancing snow cover.

stat.AP

Changepoint Detection: An Analysis of the Central England Temperature Series

This paper presents a statistical analysis of structural changes in the Central England temperature series, one of the longest surface temperature records available. A changepoint analysis is performed to detect abrupt changes, which can be regarded as a preliminary step before further analysis is conducted to identify the causes of the changes (e.g., artificial, human-induced or natural variability). Regression models with structural breaks, including mean and trend shifts, are fitted to the series and compared via two commonly used multiple changepoint penalized likelihood criteria that balance model fit quality (as measured by likelihood) against parsimony considerations. Our changepoint model fits, with independent and short-memory errors, are also compared with a different class of models termed long-memory models that have been previously used by other authors to describe persistence features in temperature series. In the end, the optimal model is judged to be one containing a changepoint in the late 1980s, with a transition to an intensified warming regime. This timing and warming conclusion is consistent across changepoint models compared in this analysis. The variability of the series is not found to be significantly changing, and shift features are judged to be more plausible than either short- or long-memory autocorrelations. The final proposed model is one including trend-shifts (both intercept and slope parameters) with independent errors. The analysis serves as a walk-through tutorial of different changepoint techniques, illustrating what can be statistically inferred.

stat.AP

A Simple Necessary Condition For Independence of Real-Valued Random Variables

The standard method to check for the independence of two real-valued random variables -- demonstrating that the bivariate joint distribution factors into the product of its marginals -- is both necessary and sufficient. Here we present a simple necessary condition based on the support sets of the random variables, which -- if not satisfied -- avoids the need to extract the marginals from the joint in demonstrating dependence. We review, in an accessible manner, the measure-theoretic, topological, and probabilistic details necessary to establish the background for the old and new ideas presented here. We prove our result in both the discrete case (where the basic ideas emerge in a simple setting), the continuous case (where serious complications emerge), and for general real-valued random variables, and we illustrate the use of our condition in three simple examples.

math.PR

Seasonal Count Time Series

Count time series are widely encountered in practice. As with continuous valued data, many count series have seasonal properties. This paper uses a recent advance in stationary count time series to develop a general seasonal count time series modeling paradigm. The model permits any marginal distribution for the series and the most flexible autocorrelations possible, including those with negative dependence. Likelihood methods of inference can be conducted and covariates can be easily accommodated. The paper first develops the modeling methods, which entail a discrete transformation of a Gaussian process having seasonal dynamics. Properties of this model class are then established and particle filtering likelihood methods of parameter estimation are developed. A simulation study demonstrating the efficacy of the methods is presented and an application to the number of rainy days in successive weeks in Seattle, Washington is given.

stat.ME

Latent Gaussian Count Time Series

This paper develops the theory and methods for modeling a stationary count time series via Gaussian transformations. The techniques use a latent Gaussian process and a distributional transformation to construct stationary series with very flexible correlation features that can have any pre-specified marginal distribution, including the classical Poisson, generalized Poisson, negative binomial, and binomial structures. Gaussian pseudo-likelihood and implied Yule-Walker estimation paradigms, based on the autocovariance function of the count series, are developed via a new Hermite expansion. Particle filtering and sequential Monte Carlo methods are used to conduct likelihood estimation. Connections to state space models are made. Our estimation approaches are evaluated in a simulation study and the methods are used to analyze a count series of weekly retail sales.

stat.ME

Autocovariance Estimation in the Presence of Changepoints

This article studies estimation of a stationary autocovariance structure in the presence of an unknown number of mean shifts. Here, a Yule-Walker moment estimator for the autoregressive parameters in a dependent time series contaminated by mean shift changepoints is proposed and studied. The estimator is based on first order differences of the series and is proven consistent and asymptotically normal when the number of changepoints $m$ and the series length $N$ satisfies $m/N \rightarrow 0$ as $N \rightarrow \infty$

math.ST

A Comparison of Single and Multiple Changepoint Techniques for Time Series Data

This paper describes and compares several prominent single and multiple changepoint techniques for time series data. Due to their importance in inferential matters, changepoint research on correlated data has accelerated recently. Unfortunately, small perturbations in model assumptions can drastically alter changepoint conclusions; for example, heavy positive correlation in a time series can be misattributed to a mean shift should correlation be ignored. This paper considers both single and multiple changepoint techniques. The paper begins by examining cumulative sum (CUSUM) and likelihood ratio tests and their variants for the single changepoint problem; here, various statistics, boundary cropping scenarios, and scaling methods (e.g., scaling to an extreme value or Brownian Bridge limit) are compared. A recently developed test based on summing squared CUSUM statistics over all times is shown to have realistic Type I errors and superior detection power. The paper then turns to the multiple changepoint setting. Here, penalized likelihoods drive the discourse, with AIC, BIC, mBIC, and MDL penalties being considered. Binary and wild binary segmentation techniques are also compared. We introduce a new distance metric specifically designed to compare two multiple changepoint segmentations. Algorithmic and computational concerns are discussed and simulations are provided to support all conclusions. In the end, the multiple changepoint setting admits no clear methodological winner, performance depending on the particular scenario. Nonetheless, some practical guidance will emerge.

stat.ME

Multiple Changepoint Detection with Partial Information on Changepoint Times

This paper proposes a new minimum description length procedure to detect multiple changepoints in time series data when some times are a priori thought more likely to be changepoints. This scenario arises with temperature time series homogenization pursuits, our focus here. Our Bayesian procedure constructs a natural prior distribution for the situation, and is shown to estimate the changepoint locations consistently, with an optimal convergence rate. Our methods substantially improve changepoint detection power when prior information is available. The methods are also tailored to bivariate data, allowing changes to occur in one or both component series.

stat.ME

A Large Scale Spatio-temporal Binomial Regression Model for Estimating Seroprevalence Trends

This paper develops a large-scale Bayesian spatio-temporal binomial regression model for the purpose of investigating regional trends in antibody prevalence to Borrelia burgdorferi, the causative agent of Lyme disease. The proposed model uses Gaussian predictive processes to estimate the spatially varying trends and a conditional autoregressive model to account for spatio-temporal dependence. Careful consideration is made to develop a novel framework that is scalable to large spatio-temporal data. The proposed model is used to analyze approximately 16 million Borrelia burgdorferi test results collected on dogs located throughout the conterminous United States over a sixty month period. This analysis identifies several regions of increasing canine risk. Specifically, this analysis reveals evidence that Lyme disease is getting worse in some endemic regions and that it could potentially be spreading to other non-endemic areas. Further, given the zoonotic nature of this vector-borne disease, this analysis could potentially reveal areas of increasing human risk.

stat.AP

An MDL approach to the climate segmentation problem

This paper proposes an information theory approach to estimate the number of changepoints and their locations in a climatic time series. A model is introduced that has an unknown number of changepoints and allows for series autocorrelations, periodic dynamics, and a mean shift at each changepoint time. An objective function gauging the number of changepoints and their locations, based on a minimum description length (MDL) information criterion, is derived. A genetic algorithm is then developed to optimize the objective function. The methods are applied in the analysis of a century of monthly temperatures from Tuscaloosa, Alabama.

stat.AP