SearcharxivSearch

arXiv subjects

Martin Tveten

Publications and source records attributed to Martin Tveten.

7 recordsLinked to original sources

skchange: Fast and Flexible Algorithms for Changepoint Detection

Skchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests. Key features include the detection of anomalous segments in addition to changepoints; theoretically well-founded fast and approximate search methods; theoretically well-founded algorithms for high-dimensional data, covering settings where either few or many features change simultaneously; utilities for automatic and data-driven penalty calibration, which balances false alarms against missed detections; and a large collection of built-in costs and statistical tests. The design follows established scikit-learn conventions to streamline both user and contributor experience, and Numba is used extensively to achieve high computational performance. Source code and documentation are available at https://github.com/NorskRegnesentral/skchange.

stat.CO

gridcp: Fast Online Changepoint Detection in Python

Online changepoint detection is the problem of detecting distributional changes in a data stream in real-time. A large body of methodology exists for the offline (fixed-size) setting, but applying these methods online quickly becomes infeasible since the per-observation computational cost and memory consumption typically grow at least linearly with the sample size. A recently proposed grid-based methodology (Moen, 2026) overcomes this by evaluating an offline test statistic over a sparse geometric grid of split points, with grid points spaced increasingly far apart further in the past. For a wide class of test statistics, this approach keeps update time and memory consumption growing logarithmic in the length of the data stream, while admitting finite-sample guarantees on the detection delay. Building on this methodology, we present gridcp, an open-source Python package that turns offline changepoint tests into efficient online detectors through a single, uniform interface. Users can choose from nine Numba-accelerated built-in tests, spanning changes in the mean, variance, covariance, and regression coefficients, as well as nonparametric tests and generalized likelihood-ratio tests for exponential-family models. Users can also supply their own test, which the package handles identically. For any test, gridcp provides Monte Carlo routines that calibrate the detection threshold to a target false alarm probability or average run length, including data-driven variants when no parametric null model is available. Through simulations and three real-data case studies, we show that calibration is accurate, that runtime scales favorably with both stream length and dimension, and that the complete pipeline, from calibration to deployment, runs efficiently on long real-world streams with only a short detection delay.

stat.ME

Fault detection in propulsion motors in the presence of concept drift

Machine learning and statistical methods can improve conventional motor protection systems, providing early warning and detection of emerging failures. Data-driven methods rely on historical data to learn how the system is expected to behave under normal circumstances. An unexpected change in the underlying system may cause a change in the statistical properties of the data, and by this alter the performance of the fault detection algorithm in terms of time to detection and false alarms. This kind of change, called \textit{concept drift}, requires adaptations to maintain constant performance. In this article, we present a machine learning approach for detecting overheating in the stator windings of marine electrical propulsion motors. Using simulated overheating faults injected into operational data, the methods are shown to provide early detection compared to conventional methods based on temperature readings and fixed limits. The proposed monitors are designed to operate for a type of concept drift observed in operational data collected from a specific class of motors in a fleet of ships. Using a mix of real and simulated concept drifts, it is shown that the proposed monitors are able to provide early detections during and after concept drifts, without the need for full model retraining.

stat.AP

Efficient sparsity adaptive changepoint estimation

We propose a new, computationally efficient, sparsity adaptive changepoint estimator for detecting changes in unknown subsets of a high-dimensional data sequence. Assuming the data sequence is Gaussian, we prove that the new method successfully estimates the number and locations of changepoints with a given error rate and under minimal conditions, for all sparsities of the changing subset. Moreover, our method has computational complexity linear up to logarithmic factors in both the length and number of time series, making it applicable to large data sets. Through extensive numerical studies we show that the new methodology is highly competitive in terms of both estimation accuracy and computational cost. The practical usefulness of the method is illustrated by analysing sensor data from a hydro power plant. An efficient R implementation is available.

stat.ME

Scalable changepoint and anomaly detection in cross-correlated data with an application to condition monitoring

Motivated by a condition monitoring application arising from subsea engineering we derive a novel, scalable approach to detecting anomalous mean structure in a subset of correlated multivariate time series. Given the need to analyse such series efficiently we explore a computationally efficient approximation of the maximum likelihood solution to the resulting modelling framework, and develop a new dynamic programming algorithm for solving the resulting Binary Quadratic Programme when the precision matrix of the time series at any given time-point is banded. Through a comprehensive simulation study, we show that the resulting methods perform favourably compared to competing methods both in the anomaly and change detection settings, even when the sparsity structure of the precision matrix estimate is misspecified. We also demonstrate its ability to correctly detect faulty time-periods of a pump within the motivating application.

stat.ME

Online Detection of Sparse Changes in High-Dimensional Data Streams Using Tailored Projections

When applying principal component analysis (PCA) for dimension reduction, the most varying projections are usually used in order to retain most of the information. For the purpose of anomaly and change detection, however, the least varying projections are often the most important ones. In this article, we present a novel method that automatically tailors the choice of projections to monitor for sparse changes in the mean and/or covariance matrix of high-dimensional data. A subset of the least varying projections is almost always selected based on a criteria of the projection's sensitivity to changes. Our focus is on online/sequential change detection, where the aim is to detect changes as quickly as possible, while controlling false alarms at a specified level. A combination of tailored PCA and a generalized log-likelihood monitoring procedure displays high efficiency in detecting even very sparse changes in the mean, variance and correlation. We demonstrate on real data that tailored PCA monitoring is efficient for sparse change detection also when the data streams are highly auto-correlated and non-normal. Notably, error control is achieved without a large validation set, which is needed in most existing methods.

stat.ME

Which principal components are most sensitive to distributional changes?

PCA is often used in anomaly detection and statistical process control tasks. For bivariate data, we prove that the minor projection (the least varying projection) of the PCA-rotated data is the most sensitive to distributional changes, where sensitivity is defined by the Hellinger distance between distributions before and after a change. In particular, this is almost always the case if only one parameter of the bivariate normal distribution changes, i.e., the change is sparse. Simulations indicate that the minor projections are the most sensitive for a large range of changes and pre-change settings in higher dimensions as well. This motivates using the minor projections for detecting sparse distributional changes in high-dimensional data.

math.ST