SearcharxivSearch

arXiv subjects

Patrick Bastian

Publications and source records attributed to Patrick Bastian.

18 recordsLinked to original sources

Change Point Detection and Localization in High-Dimensional Time Series

We present new inference tools for change point detection in high-dimensional time series. We discuss two distinct statistical applications: First, sequential change point testing in an incoming data-stream. Second, retrospective localization of multiple changes, with confidence intervals at a globally controlled error level. Test statistics are built on the maximum norm to generate power against sparse and asynchronous changes. Both problems are tackled by related multiscale statistics that search for changes in the data at many different levels of resolution. For fixed dimension, our statistical approaches can be validated using traditional H\"olderian invariance principles. In this paper, we present the high-dimensional analogue: H\"older-Gauss-approximations, which can be (roughly) interpreted as the Gaussian approximation for a H\"older-norm of the high-dimensional partial sum process. Such approximations are of interest beyond change point detection and can be used for other problems such as for stationarity testing in high dimensions. We evaluate finite-sample performance in a simulation study and give an application to air contamination due to wildfires in California, which occurs asynchronously across a panel of measuring stations.

stat.ME

Selfnormalization for relevant inference with supremum-type statistics

We develop a selfnormalized approach to inference for relevant changes in functional time series measured by the supremum norm. The main difficulty is that the supremum norm is not Hadamard differentiable, so standard projection-based selfnormalization does not apply and the limiting distribution may depend on the geometry of the extremal set and the long-run covariance structure. We address this problem by replacing the supremum norm with a smooth log-sum-exp approximation and constructing a projected selfnormalizer from its derivative. The resulting statistic has an asymptotically pivotal distribution that is free of long-run covariance nuisance parameters and depends only on the break location. We derive explicit smoothing-bias expansions for both isolated nondegenerate extrema and extremal sets of positive measure. To avoid direct estimation of geometric quantities such as the number, curvature, or measure of the extrema, we combine several smoothing levels to cancel the leading bias terms. This yields an asymptotically exact test for relevant changes under mild regularity conditions. More generally, the proposed smoothing and bias-correction principles provide a framework for combining selfnormalization with supremum-type statistics in problems involving relevant hypotheses.

math.ST

Simultaneous Inference for Partially Observed Functional Time Series

Functional data analysis (FDA) provides statistical methods for analyzing samples of time-continuous stochastic processes. Measurements often arise in the form of sensor data for a key scientific variable. The practical problem of irregular sensor disruptions has fostered interest in analyzing partially observed random functions. Specifically, this paper is motivated by a time series of intermittently missing pollution data with dependence along pollution paths and missingness patterns. To allow statistical analysis, we develop the first inference methods for dependent, partially observed functional time series. Existing methods were not appropriate for this task, because they heavily rely on the independence of the data functions. Mathematically, we model data on the space of bounded functions equipped with the supremum norm. This allows simultaneous inference across the entire functional domain, including simultaneous confidence bands -- something existing Hilbert-space-based methods cannot provide. To study non-stationary trends along the time series, we extend state-of-the-art multiscale inference methods (originally developed for scalar data) to partially observed functions. The key application of the latter methods is testing for excessive pollution levels in inner cities. Our approach combines state-of-the-art Gaussian approximations with stochastic process theory. Interestingly, it also improves existing results for fully observed functional time series by avoiding a functional CLT.

stat.ME

Differentially private testing for relevant dependencies in high dimensions

We investigate the problem of detecting dependencies between the components of a high-dimensional vector. Our approach advances the existing literature in two important respects. First, we consider the problem under privacy constraints. Second, instead of testing whether the coordinates are pairwise independent, we are interested in determining whether certain pairwise associations between the components (such as all pairwise Kendall's $\tau$ coefficients) do not exceed a given threshold in absolute value. Considering hypotheses of this form is motivated by the observation that in the high-dimensional regime, it is rare and perhaps impossible to have a null hypothesis that can be modeled exactly by assuming that all pairwise associations are precisely equal to zero. The formulation of the null hypothesis as a composite hypothesis makes the problem of constructing tests already non-standard in the non-private setting. Additionally, under privacy constraints, state of the art procedures rely on permutation approaches that are rendered invalid under a composite null. We propose a novel bootstrap based methodology that is especially powerful in sparse settings, develop theoretical guarantees under mild assumptions and show that the proposed method enjoys good finite sample properties even in the high privacy regime. Additionally, we present applications in medical data that showcase the applicability of our methodology.

math.ST

TWIN: Two window inspection for online change point detection

We propose a new class of sequential change point tests, both for changes in the mean parameter and in the overall distribution function. The methodology builds on a two-window inspection scheme (TWIN), which aggregates data into symmetric samples and applies strong weighting to enhance statistical performance. The detector yields logarithmic rather than polynomial detection delays, representing a substantial reduction compared to state-of-the-art alternatives. Delays remain short, even for late changes, where existing methods perform worst. Moreover, the new procedure also attains higher power than current methods across broad classes of local alternatives. For mean changes, we further introduce a self-normalized version of the detector that automatically cancels out temporal dependence, eliminating the need to estimate nuisance parameters. The advantages of our approach are supported by asymptotic theory, simulations and an application to monitoring COVID19 data. Here, structural breaks associated with new virus variants are detected almost immediately by our new procedures. This indicates potential value for the real-time monitoring of future epidemics. Mathematically, our approach is underpinned by new exponential moment bounds for the global modulus of continuity of the partial sum process, which may be of independent interest beyond change point testing.

math.ST

Monitoring Time Series for Relevant Changes

We consider the problem of sequentially testing for changes in the mean parameter of a time series, compared to a benchmark period. Most tests in the literature focus on the null hypothesis of a constant mean versus the alternative of a single change at an unknown time. Yet in many applications it is unrealistic that no change occurs at all, or that after one change the time series remains stationary forever. We introduce a new setup, modeling the sequence of means as a piecewise constant function with arbitrarily many changes. Instead of testing for a change, we ask whether the evolving sequence of means, say $(\mu_n)_{n \geq 1}$, stays within a narrow corridor around its initial value, that is, $\mu_n \in [\mu_1-\Delta, \mu_1+\Delta]$ for all $n \ge 1$. Combining elements from multiple change point detection with a H\"older-type monitoring procedure, we develop a new online monitoring tool. A key challenge in both construction and proof of validity is that the risk of committing a type-I error after any time $n$ fundamentally depends on the unknown future of the time series. Simulations support our theoretical results and we present two real-world applications: (1) healthcare monitoring, with a focus on blood glucose tracking, and (2) political consensus analysis via citizen opinion polls.

stat.ME

SILENT: A New Lens on Statistics in Software Timing Side Channels

Cryptographic research takes software timing side channels seriously. Approaches to mitigate them include constant-time coding and techniques to enforce such practices. However, recent attacks like Meltdown [42], Spectre [37], and Hertzbleed [70] have challenged our understanding of what it means for code to execute in constant time on modern CPUs. To ensure that assumptions on the underlying hardware are correct and to create a complete feedback loop, developers should also perform \emph{timing measurements} as a final validation step to ensure the absence of exploitable side channels. Unfortunately, as highlighted by a recent study by Jancar et al. [30], developers often avoid measurements due to the perceived unreliability of the statistical analysis and its guarantees. In this work, we combat the view that statistical techniques only provide weak guarantees by introducing a new algorithm for the analysis of timing measurements with strong, formal statistical guarantees, giving developers a reliable analysis tool. Specifically, our algorithm (1) is non-parametric, making minimal assumptions about the underlying distribution and thus overcoming limitations of classical tests like the t-test, (2) handles unknown data dependencies in measurements, (3) can estimate in advance how many samples are needed to detect a leak of a given size, and (4) allows the definition of a negligible leak threshold $\Delta$, ensuring that acceptable non-exploitable leaks do not trigger false positives, without compromising statistical soundness. We demonstrate the necessity, effectiveness, and benefits of our approach on both synthetic benchmarks and real-world applications.

cs.CR

Multiscale detection of practically significant changes in a gradually varying time series

In many change point problems it is reasonable to assume that compared to a benchmark at a given time point $t_0$ the properties of the observed stochastic process change gradually over time for $t >t_0$. Often, these gradual changes are not of interest as long as they are small (nonrelevant), but one is interested in the question if the deviations are practically significant in the sense that the deviation of the process compared to the time $t_0$ (measured by an appropriate metric) exceeds a given threshold, which is of practical significance (relevant change). In this paper we develop novel and powerful change point analysis for detecting such deviations in a sequence of gradually varying means, which is compared with the average mean from a previous time period. Current approaches to this problem suffer from low power, rely on the selection of smoothing parameters and require a rather regular (smooth) development for the means. We develop a multiscale procedure that alleviates all these issues, validate it theoretically and demonstrate its good finite sample performance on both synthetic and real data.

stat.ME

Detecting relevant dependencies under measurement error with applications to the analysis of planetary system evolution

Exoplanets play an important role in understanding the mechanics of planetary system formation and orbital evolution. In this context the correlations of different parameters of the planets and their host star are useful guides in the search for explanatory mechanisms. Based on a reanalysis of the data set from \cite{figueria14} we study the as of now still poorly understood correlation between planetary surface gravity and stellar activity of Hot Jupiters. Unfortunately, data collection often suffers from measurement errors due to complicated and indirect measurement setups, rendering standard inference techniques unreliable. We present new methods to estimate and test for correlations in a deconvolution framework and thereby improve the state of the art analysis of the data in two directions. First, we are now able to account for additive measurement errors which facilitates reliable inference. Second we test for relevant changes, i.e. we are testing for correlations exceeding a certain threshold $\Delta$. This reflects the fact that small nonzero correlations are to be expected for real life data almost always and that standard statistical tests will therefore always reject the null of no correlation given sufficient data. Our theory focuses on quantities that can be estimated by U-Statistics which contain a variety of correlation measures. We propose a bootstrap test and establish its theoretical validity. As a by product we also obtain confidence intervals. Applying our methods to the Hot Jupiter data set from \cite{figueria14}, we observe that taking into account the measurement errors yields smaller point estimates and the null of no relevant correlation is rejected only for very small $\Delta$. This demonstrates the importance of considering the impact of measurement errors to avoid misleading conclusions from the resulting statistical analysis.

stat.ME

Sequential Outlier Detection in Non-Stationary Time Series

A novel method for sequential outlier detection in non-stationary time series is proposed. The method tests the null hypothesis of ``no outlier'' at each time point, addressing the multiple testing problem by bounding the error probability of successive tests, using extreme value theory. The asymptotic properties of the test statistic are studied under the null hypothesis and alternative. The finite sample properties of the new detection scheme are investigated by means of a simulation study, and the method is compared with alternative procedures which have recently been proposed in the statistics and machine learning literature.

math.ST

Uniform confidence bands for joint angles across different fatigue phases

We develop uniform confidence bands for the mean function of stationary time series as a post-hoc analysis of multiple change point detection in functional time series. In particular, the methodology in this work provides bands for those segments where the jump size exceeds a certain threshold $\Delta$. In \cite{bastian2024multiplechangepointdetection} such exceedences of $\Delta$ were related to fatigue states of a running athlete. The extension to confidence bands stems from an interest in understanding the range of motion (ROM) of lower-extremity joints of running athletes under fatiguing conditions. From a biomechanical perspective, ROM serves as a proxy for joint flexibility under varying fatigue states, offering individualized insights into potentially problematic movement patterns. The new methodology provides a valuable tool for understanding the dynamic behavior of joint motion and its relationship to fatigue.

stat.ME

Choosing the Right Norm for Change Point Detection in Functional Data

We consider the problem of detecting a change point in a sequence of mean functions from a functional time series. We propose an $L^1$ norm based methodology and establish its theoretical validity both for classical and for relevant hypotheses. We compare the proposed method with currently available methodology that is based on the $L^2$ and supremum norms. Additionally we investigate the asymptotic behaviour under the alternative for all three methods and showcase both theoretically and empirically that the $L^1$ norm achieves the best performance in a broad range of scenarios. We also propose a power enhancement component that improves the performance of the $L^1$ test against sparse alternatives. Finally we apply the proposed methodology to both synthetic and real data.

math.ST

Detecting relevant deviations from the white noise assumption for non-stationary time series

We consider the problem of detecting deviations from a white noise assumption in time series. Our approach differs from the numerous methods proposed for this purpose with respect to two aspects. First, we allow for non-stationary time series. Second, we address the problem that a white noise test, for example checking the residuals of a model fit, is usually not performed because one believes in this hypothesis, but thinks that the white noise hypothesis may be approximately true, because a postulated models describes the unknown relation well. This reflects a meanwhile classical paradigm of Box(1976) that "all models are wrong but some are useful". We address this point of view by investigating if the maximum deviation of the local autocovariance functions from 0 exceeds a given threshold $\Delta$ that can either be specified by the user or chosen in a data dependent way. The formulation of the problem in this form raises several mathematical challenges, which do not appear when one is testing the classical white noise hypothesis. We use high dimensional Gaussian approximations for dependent data to furnish a bootstrap test, prove its validity and showcase its performance on both synthetic and real data, in particular we inspect log returns of stock prices and show that our approach reflects some observations of Fama(1970) regarding the efficient market hypothesis.

math.ST

Gradual changes in functional time series

We consider the problem of detecting gradual changes in the sequence of mean functions from a not necessarily stationary functional time series. Our approach is based on the maximum deviation (calculated over a given time interval) between a benchmark function and the mean functions at different time points. We speak of a gradual change of size $\Delta $, if this quantity exceeds a given threshold $\Delta>0$. For example, the benchmark function could represent an average of yearly temperature curves from the pre-industrial time, and we are interested in the question if the yearly temperature curves afterwards deviate from the pre-industrial average by more than $\Delta =1.5$ degrees Celsius, where the deviations are measured with respect to the sup-norm. Using Gaussian approximations for high-dimensional data we develop a test for hypotheses of this type and estimators for the time where a deviation of size larger than $\Delta$ appears for the first time. We prove the validity of our approach and illustrate the new methods by a simulation study and a data example, where we analyze yearly temperature curves at different stations in Australia.

math.ST

Multiple change point detection in functional data with applications to biomechanical fatigue data

Injuries to the lower extremity joints are often debilitating, particularly for professional athletes. Understanding the onset of stressful conditions on these joints is therefore important in order to ensure prevention of injuries as well as individualised training for enhanced athletic performance. We study the biomechanical joint angles from the hip, knee and ankle for runners who are experiencing fatigue. The data is cyclic in nature and densely collected by body worn sensors, which makes it ideal to work with in the functional data analysis (FDA) framework. We develop a new method for multiple change point detection for functional data, which improves the state of the art with respect to at least two novel aspects. First, the curves are compared with respect to their maximum absolute deviation, which leads to a better interpretation of local changes in the functional data compared to classical $L^2$-approaches. Secondly, as slight aberrations are to be often expected in a human movement data, our method will not detect arbitrarily small changes but hunts for relevant changes, where maximum absolute deviation between the curves exceeds a specified threshold, say $\Delta >0$. We recover multiple changes in a long functional time series of biomechanical knee angle data, which are larger than the desired threshold $\Delta$, allowing us to identify changes purely due to fatigue. In this work, we analyse data from both controlled indoor as well as from an uncontrolled outdoor (marathon) setting.

math.ST

Testing equivalence of multinomial distributions -- a constrained bootstrap approach

In this paper we develop a novel bootstrap test for the comparison of two multinomial distributions. The two distributions are called {\it equivalent} or {\it similar} if a norm of the difference between the class probabilities is smaller than a given threshold. In contrast to most of the literature our approach does not require differentiability of the norm and is in particular applicable for the maximum- and $L^1$-norm.

math.ST

Comparing regression curves -- an $L^1$-point of view

In this paper we compare two regression curves by measuring their difference by the area between the two curves, represented by their $L^1$-distance. We develop asymptotic confidence intervals for this measure and statistical tests to investigate the similarity/equivalence of the two curves. Bootstrap methodology specifically designed for equivalence testing is developed to obtain procedures with good finite sample properties and its consistency is rigorously proved. The finite sample properties are investigated by means of a small simulation study.

math.ST

Testing for practically significant dependencies in high dimensions via bootstrapping maxima of U-statistics

This paper takes a different look on the problem of testing the mutual independence of the components of a high-dimensional vector. Instead of testing if all pairwise associations (e.g. all pairwise Kendall's $\tau$) between the components vanish, we are interested in the (null)-hypothesis that all pairwise associations do not exceed a certain threshold in absolute value. The consideration of these hypotheses is motivated by the observation that in the high-dimensional regime, it is rare, and perhaps impossible, to have a null hypothesis that can be exactly modeled by assuming that all pairwise associations are precisely equal to zero. The formulation of the null hypothesis as a composite hypothesis makes the problem of constructing tests non-standard and in this paper we provide a solution for a broad class of dependence measures, which can be estimated by $U$-statistics. In particular we develop an asymptotic and a bootstrap level $\alpha$-test for the new hypotheses in the high-dimensional regime. We also prove that the new tests are minimax-optimal and investigate their finite sample properties by means of a small simulation study and a data example.

math.ST