SearcharxivSearch

arXiv subjects

Rebecca Killick

Publications and source records attributed to Rebecca Killick.

At least 19 recordsLinked to original sources

High Dimensional Change Point Models for Two-Directional Data

We develop methodology for recovery of change points for data observed on more than one temporal index where changes may occur simultaneous in both indices, where the spatial component may be high dimensional. The work is motivated by climate monitoring problems where long series of data are available, e.g., daily observations (index 1) over several years (index 2). Such data may be evolving over the annual time scale, along with dynamic seasonal changes in the shorter time scale. We model this as a high dimensional mean process observed on a two dimensional grid with change points. Asymptotic estimation and inference results are developed under a single change point setup, including rates of convergence of the proposed method as well the resulting limiting distributions. The method is extended to the case of multiple changes. Theoretical results are supported numerically with monte-carlo simulations. We implement our work on a large scale climate data for the Pacific Northwest region of the United States.

stat.ME

An ensemble prediction method for forecasting sap flux density and water-use in temperate trees

Efficient irrigation management is crucial to agriculture, forestry and horticulture, especially under climate change. Developments in novel sensors and Internet of Things technology provide an opportunity to carry out real-time monitoring of tree sap flux density, which, when coupled with advanced modelling techniques, enables online prediction of tree water-use suitable for irrigation planning. This manuscript proposes one such pipeline that integrates tree sap flow sensors, weather station sensors, and statistical models to predict tree daily water-use. In particular, an ensemble prediction approach based on additive models has been developed, using weather data as the main predictors of sap flux density. The method simultaneously considers the non-linear relationships and interactions between sap flux density and its environmental drivers, as well as the variability among individual trees over different growing seasons. Using field data collected on nine species of trees over the 2022, 2023 and 2024 growing seasons, this manuscript demonstrates the ability of the proposed ensemble prediction method in producing reliable daily water-use forecasts. The challenge of predicting tree water-use under climate stress, such as heatwaves, and the impact of tree sizes on prediction have also been discussed. Despite the complexity of the problem, the proposed method provides a general framework which can be used in a variety of settings, from commercial tree growers to conversation work. The model can be integrated into an online monitoring platform, assisting real-time decision making on irrigation management.

stat.AP

Overcoming Barriers to Computational Reproducibility

Computational reproducibility, the possibility for independent researchers to exactly reproduce published empirical results, is fundamental to science. Despite its importance, the proportion of research articles aiming for reproducibility remains low and uneven across disciplines. Barriers include a perceived lack of incentives for researchers and journals, practical challenges in preparing reproducible materials, and the absence of harmonised standards of reproducibility processes and requirements by journals. Existing guidance is often highly technical, reaching mainly those already engaged with reproducible research. In this paper, we first synthesize evidence on the benefits of reproducibility for both authors and journals. Drawing on our extensive experience in reproducibility checking at various journals, we then put forward concise, pragmatic guidelines for creating reproducible analyses across disciplines. We further review current reproducibility policies of selected journals, illustrating the substantial heterogeneity in requirements and procedures. Motivated by the latter, we propose conceptual foundations for a harmonised multi-tier system of reproducibility standards that could support transparent, consistent assessment across journals and research communities. Our goal as journal (reproducibility) editors and contributors to the MaRDI initiative is to encourage broader adoption of reproducibility practices, in particular by lowering practical barriers for authors and journals.

cs.DL

Decoder-only Clustering in Attributed Graphs

This manuscript studies nodal clustering in graphs having multivariate attributes at each node. The framework includes node-specific priors for low-dimensional representations, coupled with a neural decoder that bridges observed attributes with latent variables. Structural and attribute information are incorporated through a graph-fused LASSO regularization on the prior means, promoting nodal clustering. The optimization problem is solved via alternating direction method of multipliers, with Langevin dynamics for posterior inference. Simulation studies on grid graphs, and applications to real data with complex settings, demonstrate the effectiveness of the proposed clustering method.

stat.ME

Inferring Soil Drydown Behaviour with Adaptive Bayesian Online Changepoint Analysis

Continuous soil-moisture measurements provide a direct lens on subsurface hydrological processes, notably the post-rainfall "drydown" phase. Because these records consist of distinct, segment-specific behaviours whose forms and scales vary over time, realistic inference demands a model that captures piecewise dynamics while accommodating parameters that are unknown a priori. Building on Bayesian Online Changepoint Detection (BOCPD), we introduce two complementary extensions: a particle-filter variant that substitutes exact marginalisation with sequential Monte Carlo to enable real-time inference when critical parameters cannot be integrated out analytically, and an online-gradient variant that embeds stochastic gradient updates within BOCPD to learn application-relevant parameters on the fly without prohibitive computational cost. After validating both algorithms on synthetic data that replicate the temporal structure of field observations-detailing hyperparameter choices, priors, and cost-saving strategies-we apply them to soil-moisture series from experimental sites in Austria and the United States, quantifying site-specific drydown rates and demonstrating the advantages of our adaptive framework over static models.

stat.AP

Online detection of forecast model inadequacies using forecast errors

In many organisations, accurate forecasts are essential for making informed decisions for a variety of applications from inventory management to staffing optimization. Whatever forecasting model is used, changes in the underlying process can lead to inaccurate forecasts, which will be damaging to decision-making. At the same time, models are becoming increasingly complex and identifying change through direct modelling is problematic. We present a novel framework for online monitoring of forecasts to ensure they remain accurate. By utilizing sequential changepoint techniques on the forecast errors, our framework allows for the real-time identification of potential changes in the process caused by various external factors. We show theoretically that some common changes in the underlying process will manifest in the forecast errors and can be identified faster by identifying shifts in the forecast errors than within the original modelling framework. Moreover, we demonstrate the effectiveness of this framework on numerous forecasting approaches through simulations and show its effectiveness over alternative approaches. Finally, we present two concrete examples, one from Royal Mail parcel delivery volumes and one from NHS A\&E admissions relating to gallstones.

stat.ME

Generalised mixed effects models for changepoint analysis of biomedical time series data

Motivated by two distinct types of biomedical time series data, digital health monitoring and neuroimaging, we develop a novel approach for changepoint analysis that uses a generalised linear mixed model framework. The generalised linear mixed model framework lets us incorporate structure that is usually present in biomedical time series data. We embed the mixed model in a dynamic programming algorithm for detecting multiple changepoints in the fMRI data. We evaluate the performance of our proposed method across several scenarios using simulations. Finally, we show the utility of our proposed method on our two distinct motivating applications.

stat.ME

TrendLSW: Trend and Spectral Estimation of Nonstationary Time Series in R

The TrendLSW R package has been developed to provide users with a suite of wavelet-based techniques to analyse the statistical properties of nonstationary time series. The key components of the package are (a) two approaches for the estimation of the evolutionary wavelet spectrum in the presence of trend; and (b) wavelet-based trend estimation in the presence of locally stationary wavelet errors via both linear and nonlinear wavelet thresholding; and (c) the calculation of associated pointwise confidence intervals. Lastly, the package directly implements boundary handling methods that enable the methods to be performed on data of arbitrary length, not just dyadic length as is common for wavelet-based methods, ensuring no pre-processing of data is necessary. The key functionality of the package is demonstrated through two data examples, arising from biology and activity monitoring.

stat.ME

Is a Recent Surge in Global Warming Detectable?

The global mean surface temperature is widely studied to monitor climate change. A current debate centers around whether there has been a recent (post-1970s) surge/acceleration in the warming rate. This paper addresses whether an acceleration in the warming rate is detectable from a statistical perspective. We use changepoint models, which are statistical techniques specifically designed for identifying structural changes in time series. Four global mean surface temperature records over 1850-2023 are scrutinized within. Our results show limited evidence for a warming surge; in most surface temperature time series, no change in the warming rate beyond the 1970s is detected. As such, we estimate minimum changes in the warming trend for a surge to be detectable in the near future.

stat.AP

Statistical monitoring of European cross-border physical electricity flows using novel temporal edge network processes

Conventional modelling of networks evolving in time focuses on capturing variations in the network structure. However, the network might be static from the origin or experience only deterministic, regulated changes in its structure, providing either a physical infrastructure or a specified connection arrangement for some other processes. Thus, to detect change in its exploitation, we need to focus on the processes happening on the network. In this work, we present the concept of monitoring random Temporal Edge Network (TEN) processes that take place on the edges of a graph having a fixed structure. Our framework is based on the Generalized Network Autoregressive statistical models with time-dependent exogenous variables (GNARX models) and Cumulative Sum (CUSUM) control charts. To demonstrate its effective detection of various types of change, we conduct a simulation study and monitor the real-world data of cross-border physical electricity flows in Europe.

stat.AP

A changepoint approach to modelling non-stationary soil moisture dynamics

Soil moisture dynamics provide an indicator of soil health that scientists model via drydown curves. The typical modelling process requires the soil moisture time series to be manually separated into drydown segments and then exponential decay models are fitted to them independently. Sensor development over recent years means that experiments that were previously conducted over a few field campaigns can now be scaled to months or years at a higher sampling rate. To better meet the challenge of increasing data size, this paper proposes a novel changepoint-based approach to automatically identify structural changes in the soil drying process and simultaneously estimate the drydown parameters that are of interest to soil scientists. A simulation study is carried out to demonstrate the performance of the method in detecting changes and retrieving model parameters. Practical aspects of the method such as adding covariates and penalty learning are discussed. The method is applied to hourly soil moisture time series from the NEON data portal to investigate the temporal dynamics of soil moisture drydown. We recover known relationships previously identified manually, alongside delivering new insights into the temporal variability across soil types and locations.

stat.AP

Automatic Locally Stationary Time Series Forecasting with application to predicting U.K. Gross Value Added Time Series under sudden shocks caused by the COVID pandemic

Accurate forecasting of the U.K. gross value added (GVA) is fundamental for measuring the growth of the U.K. economy. A common nonstationarity in GVA data, such as the ABML series, is its increase in variance over time due to inflation. Transformed or inflation-adjusted series can still be challenging for classical stationarity-assuming forecasters. We adopt a different approach that works directly with the GVA series by advancing recent forecasting methods for locally stationary time series. Our approach results in more accurate and reliable forecasts, and continues to work well even when the ABML series becomes highly variable during the COVID pandemic.

stat.ME

Good Practices and Common Pitfalls in Climate Time Series Changepoint Techniques: A Review

Climate changepoint (homogenization) methods abound today, with a myriad of techniques existing in both the climate and statistics literature. Unfortunately, the appropriate changepoint technique to use remains unclear to many. Further complicating issues, changepoint conclusions are not robust to small perturbations in assumptions; for example, allowing for a trend or correlation in the series can drastically change conclusions. This paper is a review of the changepoint topic, with an emphasis on illuminating the models and techniques that allow the scientist to make reliable conclusions. Pitfalls to avoid are demonstrated via actual applications. The discourse begins by narrating the salient statistical features of most climate time series. Thereafter, single and multiple changepoint problems are considered. Several pitfalls are discussed en route and good practices are recommended. While the majority of our applications involve temperature series, other settings are mentioned.

stat.AP

Changepoint Detection: An Analysis of the Central England Temperature Series

This paper presents a statistical analysis of structural changes in the Central England temperature series, one of the longest surface temperature records available. A changepoint analysis is performed to detect abrupt changes, which can be regarded as a preliminary step before further analysis is conducted to identify the causes of the changes (e.g., artificial, human-induced or natural variability). Regression models with structural breaks, including mean and trend shifts, are fitted to the series and compared via two commonly used multiple changepoint penalized likelihood criteria that balance model fit quality (as measured by likelihood) against parsimony considerations. Our changepoint model fits, with independent and short-memory errors, are also compared with a different class of models termed long-memory models that have been previously used by other authors to describe persistence features in temperature series. In the end, the optimal model is judged to be one containing a changepoint in the late 1980s, with a transition to an intensified warming regime. This timing and warming conclusion is consistent across changepoint models compared in this analysis. The variability of the series is not found to be significantly changing, and shift features are judged to be more plausible than either short- or long-memory autocorrelations. The final proposed model is one including trend-shifts (both intercept and slope parameters) with independent errors. The analysis serves as a walk-through tutorial of different changepoint techniques, illustrating what can be statistically inferred.

stat.AP

Modelling Time-Varying First and Second-Order Structure of Time Series via Wavelets and Differencing

Most time series observed in practice exhibit time-varying trend (first-order) and autocovariance (second-order) behaviour. Differencing is a commonly-used technique to remove the trend in such series, in order to estimate the time-varying second-order structure (of the differenced series). However, often we require inference on the second-order behaviour of the original series, for example, when performing trend estimation. In this article, we propose a method, using differencing, to jointly estimate the time-varying trend and second-order structure of a nonstationary time series, within the locally stationary wavelet modelling framework. We develop a wavelet-based estimator of the second-order structure of the original time series based on the differenced estimate, and show how this can be incorporated into the estimation of the trend of the time series. We perform a simulation study to investigate the performance of the methodology, and demonstrate the utility of the method by analysing data examples from environmental and biomedical science.

stat.ME

Detecting changes in covariance via random matrix theory

A novel method is proposed for detecting changes in the covariance structure of moderate dimensional time series. This non-linear test statistic has a number of useful properties. Most importantly, it is independent of the underlying structure of the covariance matrix. We discuss how results from Random Matrix Theory, can be used to study the behaviour of our test statistic in a moderate dimensional setting (i.e. the number of variables is comparable to the length of the data). In particular, we demonstrate that the test statistic converges point wise to a normal distribution under the null hypothesis. We evaluate the performance of the proposed approach on a range of simulated datasets and find that it outperforms a range of alternative recently proposed methods. Finally, we use our approach to study changes in the amount of water on the surface of a plot of soil which feeds into model development for degradation of surface piping.

stat.ME

Graphical Influence Diagnostics for Changepoint Models

Changepoint models enjoy a wide appeal in a variety of disciplines to model the heterogeneity of ordered data. Graphical influence diagnostics to characterize the influence of single observations on changepoint models are, however, lacking. We address this gap by developing a framework for investigating instabilities in changepoint segmentations and assessing the influence of single observations on various outputs of a changepoint analysis. We construct graphical diagnostic plots that allow practitioners to assess whether instabilities occur; how and where they occur; and to detect influential individual observations triggering instability. We analyze well-log data to illustrate how such influence diagnostic plots can be used in practice to reveal features of the data that may otherwise remain hidden.

stat.ME

Autocovariance Estimation in the Presence of Changepoints

This article studies estimation of a stationary autocovariance structure in the presence of an unknown number of mean shifts. Here, a Yule-Walker moment estimator for the autoregressive parameters in a dependent time series contaminated by mean shift changepoints is proposed and studied. The estimator is based on first order differences of the series and is proven consistent and asymptotically normal when the number of changepoints $m$ and the series length $N$ satisfies $m/N \rightarrow 0$ as $N \rightarrow \infty$

math.ST