SearcharxivSearch

arXiv subjects

Lorenzo Trapani

Publications and source records attributed to Lorenzo Trapani.

At least 19 recordsLinked to original sources

Online detection of distributional changes for time series in metric spaces

We propose an online testing framework for detecting distributional changes in serially dependent data with values in a separable metric space. Based on two-sample $U$-statistics, the framework encompasses sequential analogs of energy distance and maximum mean discrepancy (MMD) procedures while accommodating temporal dependence. We establish asymptotic theory for finite and open-ended monitoring horizons that characterizes the full asymptotic run-length distribution under $H_0$ and yields asymptotic false-alarm control. We further establish new spectral approximation results for kernel matrices formed from serially dependent observations, and use them to construct a feasible Monte Carlo calibration procedure. Our flexible window construction encompasses classical, Page-type, and full-scan historical-baseline monitoring and can achieve short detection delays for both early and late changepoints, without requiring sub-Gaussianity or high-order moments of the raw observations. Simulations show reliable false-alarm control across linear, nonlinear, high-dimensional, and functional time-series models and further demonstrate that, over a broad range of alternatives and changepoint locations, the proposed method can achieve substantially shorter delays than recent procedures designed specifically for rapid detection. Applications to foreign exchange rates, electricity-market curves, and daily air transportation networks illustrate the methodology across scalar, functional, and network-valued time series.

stat.ME

Sequential monitoring for distributional changepoints using degenerate U-statistics

We investigate the online detection of changepoints in the distribution of a sequence of observations using a class of degenerate \textit{U}-statistic-type processes. We consider an ordinary (Chu--Stinchcombe--White-type) detector and a Page-type detector under open- and closed-ended monitoring, and introduce an expanding-baseline Page-type procedure that incorporates sufficiently old monitoring observations into the baseline sample. Under the null, we derive weak limits for all three procedures and justify a Monte Carlo approximation to their critical values. For the ordinary and Page-type detectors, we also establish consistency and limiting distributions for detection delays under both early and late changes. The theory requires only square summability of the eigenvalues associated with the degenerate kernel operator, rather than the stronger absolute-summability condition often imposed in related work. Simulations show competitive performance relative to recent mean-, covariance-, and empirical-CDF-based monitors, and an application to multivariate compressor-sensor data from a metro train illustrates the methodology.

math.ST

A general randomized test for Alpha

We propose a methodology to construct tests for the null hypothesis that the pricing errors of a panel of asset returns are jointly equal to zero in a linear factor asset pricing model -- that is, the null of "zero alpha". We consider, as a leading example, a model with observable, tradable factors, but we also develop extensions to accommodate for non-tradable and latent factors. The test is based on equation-by-equation estimation, using a randomized version of the estimated alphas, which only requires rates of convergence. The distinct features of the proposed methodology are that it does not require the estimation of any covariance matrix, and that it allows for both N and T to pass to infinity, with the former possibly faster than the latter. Further, unlike extant approaches, the procedure can accommodate conditional heteroskedasticity, non-Gaussianity, and even strong cross-sectional dependence in the error terms. We also propose a de-randomized decision rule to choose in favor or against the correct specification of a linear factor pricing model. Monte Carlo simulations show that the test has satisfactory properties and it compares favorably to several existing tests. The usefulness of the testing procedure is illustrated through an application of linear factor pricing models to price the constituents of the S&P 500.

econ.EM

Moving sum procedure for multiple change point detection in large factor models

This paper proposes a moving sum methodology for detecting multiple change points in high-dimensional time series under a factor model, where changes are attributed to those in loadings as well as emergence or disappearance of factors. We establish the asymptotic null distribution of the proposed test for family-wise error control, and show the consistency of the procedure for multiple change point estimation. Simulation studies and an application to a large dataset of volatilities demonstrate the competitive performance of the proposed method.

stat.ME

Statistical inference for large-dimensional tensor factor model by iterative projections

Tensor Factor Models (TFM) are appealing dimension reduction tools for high-order large-dimensional tensor time series, and have wide applications in economics, finance and medical imaging. In this paper, we propose a projection estimator for the Tucker-decomposition based TFM, and provide its least-square interpretation which parallels to the least-square interpretation of the Principal Component Analysis (PCA) for the vector factor model. The projection technique simultaneously reduces the dimensionality of the signal component and the magnitudes of the idiosyncratic component tensor, thus leading to an increase of the signal-to-noise ratio. We derive a convergence rate of the projection estimator of the loadings and the common factor tensor which are faster than that of the naive PCA-based estimator. Our results are obtained under mild conditions which allow the idiosyncratic components to be weakly cross- and auto- correlated. We also provide a novel iterative procedure based on the eigenvalue-ratio principle to determine the factor numbers. Extensive numerical studies are conducted to investigate the empirical performance of the proposed projection estimators relative to the state-of-the-art ones.

stat.ME

Inference in matrix-valued time series with common stochastic trends and multifactor error structure

We develop an estimation methodology for a factor model for high-dimensional matrix-valued time series, where common stochastic trends and common stationary factors can be present. We study, in particular, the estimation of (row and column) loading spaces, of the common stochastic trends and of the common stationary factors, and the row and column ranks thereof. In a set of (negative) preliminary results, we show that a projection-based technique fails to improve the rates of convergence compared to a "flattened" estimation technique which does not take into account the matrix nature of the data. Hence, we develop a three-step algorithm where: (i) we first project the data onto the orthogonal complement to the (row and column) loadings of the common stochastic trends; (ii) we subsequently use such "trend free" data to estimate the stationary common component; (iii) we remove the estimated common stationary component from the data, and re-estimate, using a projection-based estimator, the row and column common stochastic trends and their loadings. We show that this estimator succeeds in refining the rates of convergence of the initial, "flattened" estimator. As a by-product, we develop consistent eigenvalue-ratio based estimators for the number of stationary and nonstationary common factors.

stat.ME

Sequential monitoring for explosive volatility regimes

In this paper, we develop two families of sequential monitoring procedure to (timely) detect changes in a GARCH(1,1) model. Whilst our methodologies can be applied for the general analysis of changepoints in GARCH(1,1) sequences, they are in particular designed to detect changes from stationarity to explosivity or vice versa, thus allowing to check for volatility bubbles. Our statistics can be applied irrespective of whether the historical sample is stationary or not, and indeed without prior knowledge of the regime of the observations before and after the break. In particular, we construct our detectors as the CUSUM process of the quasi-Fisher scores of the log likelihood function. In order to ensure timely detection, we then construct our boundary function (exceeding which would indicate a break) by including a weighting sequence which is designed to shorten the detection delay in the presence of a changepoint. We consider two types of weights: a lighter set of weights, which ensures timely detection in the presence of changes occurring early, but not too early after the end of the historical sample; and a heavier set of weights, called Renyi weights which is designed to ensure timely detection in the presence of changepoints occurring very early in the monitoring horizon. In both cases, we derive the limiting distribution of the detection delays, indicating the expected delay for each set of weights. Our theoretical results are validated via a comprehensive set of simulations, and an empirical application to daily returns of individual stocks.

econ.EM

Fast Online Changepoint Detection

We study online changepoint detection in the context of a linear regression model. We propose a class of heavily weighted statistics based on the CUSUM process of the regression residuals, which are specifically designed to ensure timely detection of breaks occurring early on during the monitoring horizon. We subsequently propose a class of composite statistics, constructed using different weighing schemes; the decision rule to mark a changepoint is based on the largest statistic across the various weights, thus effectively working like a veto-based voting mechanism, which ensures fast detection irrespective of the location of the changepoint. Our theory is derived under a very general form of weak dependence, thus being able to apply our tests to virtually all time series encountered in economics, medicine, and other applied sciences. Monte Carlo simulations show that our methodologies are able to control the procedure-wise Type I Error, and have short detection delays in the presence of breaks.

stat.ME

Real-time monitoring with RCA models

We propose a family of weighted statistics based on the CUSUM process of the WLS residuals for the online detection of changepoints in a Random Coefficient Autoregressive model, using both the standard CUSUM and the Page-CUSUM process. We derive the asymptotics under the null of no changepoint for all possible weighing schemes, including the case of the standardised CUSUM, for which we derive a Darling-Erdos-type limit theorem; our results guarantee the procedure-wise size control under both an open-ended and a closed-ended monitoring. In addition to considering the standard RCA model with no covariates, we also extend our results to the case of exogenous regressors. Our results can be applied irrespective of (and with no prior knowledge required as to) whether the observations are stationary or not, and irrespective of whether they change into a stationary or nonstationary regime. Hence, our methodology is particularly suited to detect the onset, or the collapse, of a bubble or an epidemic. Our simulations show that our procedures, especially when standardising the CUSUM process, can ensure very good size control and short detection delays. We complement our theory by studying the online detection of breaks in epidemiological and housing prices series.

stat.ME

On changepoint detection in functional data using empirical energy distance

We propose a novel family of test statistics to detect the presence of changepoints in a sequence of dependent, possibly multivariate, functional-valued observations. Our approach allows to test for a very general class of changepoints, including the "classical" case of changes in the mean, and even changes in the whole distribution. Our statistics are based on a generalisation of the empirical energy distance; we propose weighted functionals of the energy distance process, which are designed in order to enhance the ability to detect breaks occurring at sample endpoints. The limiting distribution of the maximally selected version of our statistics requires only the computation of the eigenvalues of the covariance function, thus being readily implementable in the most commonly employed packages, e.g. R. We show that, under the alternative, our statistics are able to detect changepoints occurring even very close to the beginning/end of the sample. In the presence of multiple changepoints, we propose a binary segmentation algorithm to estimate the number of breaks and the locations thereof. Simulations show that our procedures work very well in finite samples. We complement our theory with applications to financial and temperature data.

stat.ME

Robust Tensor Factor Analysis

We consider (robust) inference in the context of a factor model for tensor-valued sequences. We study the consistency of the estimated common factors and loadings space when using estimators based on minimising quadratic loss functions. Building on the observation that such loss functions are adequate only if sufficiently many moments exist, we extend our results to the case of heavy-tailed distributions by considering estimators based on minimising the Huber loss function, which uses an $L_{1}$ -norm weight on outliers. We show that such class of estimators is robust to the presence of heavy tails, even when only the second moment of the data exists. We also propose a modified version of the eigenvalue-ratio principle to estimate the dimensions of the core tensor and show the consistency of the resultant estimators without any condition on the relative rates of divergence of the sample size and dimensions. Extensive numerical studies are conducted to show the advantages of the proposed methods over the state-of-the-art ones especially under the heavy-tailed cases. An import/export dataset of a variety of commodities across multiple countries is analyzed to show the practical usefulness of the proposed robust estimation procedure. An R package ``RTFA" implementing the proposed methods is available on R CRAN.

stat.ME

One-way or Two-way Factor Model for Matrix Sequences?

This paper investigates the issue of determining the dimensions of row and column factor spaces in matrix-valued data. Exploiting the eigen-gap in the spectrum of sample second moment matrices of the data, we propose a family of randomised tests to check whether a one-way or two-way factor structure exists or not. Our tests do not require any arbitrary thresholding on the eigenvalues, and can be applied with no restrictions on the relative rate of divergence of the cross-sections to the sample sizes as they pass to infinity. Although tests are based on a randomization which does not vanish asymptotically, we propose a de-randomized, strong (based on the Law of the Iterated Logarithm) decision rule to choose in favor or against the presence of common factors. We use the proposed tests and decision rule in two ways. We further cast our individual tests in a sequential procedure whose output is an estimate of the number of common factors. Our tests are built on two variants of the sample second moment matrix of the data: one based on a row (or column) flattened version of the matrix-valued sequence, and one based on a projection-based method. Our simulations show that both procedures work well in large samples and, in small samples, the one based on the projection method delivers a superior performance compared to existing methods in virtually all cases considered.

stat.ME

Online Change-point Detection for Matrix-valued Time Series with Latent Two-way Factor Structure

This paper proposes a novel methodology for the online detection of changepoints in the factor structure of large matrix time series. Our approach is based on the well-known fact that, in the presence of a changepoint, a factor model can be rewritten as a model with a larger number of common factors. In turn, this entails that, in the presence of a changepoint, the number of spiked eigenvalues in the second moment matrix of the data increases. Based on this, we propose two families of procedures - one based on the fluctuations of partial sums, and one based on extreme value theory - to monitor whether the first non-spiked eigenvalue diverges after a point in time in the monitoring horizon, thereby indicating the presence of a changepoint. Our procedure is based only on rates; at each point in time, we randomise the estimated eigenvalue, thus obtaining a normally distributed sequence which is $i.i.d.$ with mean zero under the null of no break, whereas it diverges to positive infinity in the presence of a changepoint. We base our monitoring procedures on such sequence. Extensive simulation studies and empirical analysis justify the theory.

stat.ME

Inference in heavy-tailed non-stationary multivariate time series

We study inference on the common stochastic trends in a non-stationary, $N$-variate time series $y_{t}$, in the possible presence of heavy tails. We propose a novel methodology which does not require any knowledge or estimation of the tail index, or even knowledge as to whether certain moments (such as the variance) exist or not, and develop an estimator of the number of stochastic trends $m$ based on the eigenvalues of the sample second moment matrix of $y_{t}$. We study the rates of such eigenvalues, showing that the first $m$ ones diverge, as the sample size $T$ passes to infinity, at a rate faster by $O\left(T \right)$ than the remaining $N-m$ ones, irrespective of the tail index. We thus exploit this eigen-gap by constructing, for each eigenvalue, a test statistic which diverges to positive infinity or drifts to zero according to whether the relevant eigenvalue belongs to the set of the first $m$ eigenvalues or not. We then construct a randomised statistic based on this, using it as part of a sequential testing procedure, ensuring consistency of the resulting estimator of $m$. We also discuss an estimator of the common trends based on principal components and show that, up to a an invertible linear transformation, such estimator is consistent in the sense that the estimation error is of smaller order than the trend itself. Finally, we also consider the case in which we relax the standard assumption of \textit{i.i.d.} innovations, by allowing for heterogeneity of a very general form in the scale of the innovations. A Monte Carlo study shows that the proposed estimator for $m$ performs particularly well, even in samples of small size. We complete the paper by presenting four illustrative applications covering commodity prices, interest rates data, long run PPP and cryptocurrency markets.

econ.EM

Changepoint detection in random coefficient autoregressive models

We propose a family of CUSUM-based statistics to detect the presence of changepoints in the deterministic part of the autoregressive parameter in a Random Coefficient AutoRegressive (RCA) sequence. In order to ensure the ability to detect breaks at sample endpoints, we thoroughly study weighted CUSUM statistics, analysing the asymptotics for virtually all possible weighing schemes, including the standardised CUSUM process (for which we derive a Darling-Erdos theorem) and even heavier weights (studying the so-called Rényi statistics). Our results are valid irrespective of whether the sequence is stationary or not, and no prior knowledge of stationarity or lack thereof is required. Technically, our results require strong approximations which, in the nonstationary case, are entirely new. Similarly, we allow for heteroskedasticity of unknown form in both the error term and in the stochastic part of the autoregressive coefficient, proposing a family of test statistics which are robust to heteroskedasticity, without requiring any prior knowledge as to the presence or type thereof. Simulations show that our procedures work very well in finite samples. We complement our theory with applications to financial, economic and epidemiological time series.

math.ST

Sequential monitoring for cointegrating regressions

We develop monitoring procedures for cointegrating regressions, testing the null of no breaks against the alternatives that there is either a change in the slope, or a change to non-cointegration. After observing the regression for a calibration sample m, we study a CUSUM-type statistic to detect the presence of change during a monitoring horizon m+1,...,T. Our procedures use a class of boundary functions which depend on a parameter whose value affects the delay in detecting the possible break. Technically, these procedures are based on almost sure limiting theorems whose derivation is not straightforward. We therefore define a monitoring function which - at every point in time - diverges to infinity under the null, and drifts to zero under alternatives. We cast this sequence in a randomised procedure to construct an i.i.d. sequence, which we then employ to define the detector function. Our monitoring procedure rejects the null of no break (when correct) with a small probability, whilst it rejects with probability one over the monitoring horizon in the presence of breaks.

econ.EM

Sequential testing for structural stability in approximate factor models

We develop a monitoring procedure to detect changes in a large approximate factor model. Letting $r$ be the number of common factors, we base our statistics on the fact that the $\left( r+1\right) $-th eigenvalue of the sample covariance matrix is bounded under the null of no change, whereas it becomes spiked under changes. Given that sample eigenvalues cannot be estimated consistently under the null, we randomise the test statistic, obtaining a sequence of \textit{i.i.d} statistics, which are used for the monitoring scheme. Numerical evidence shows a very small probability of false detections, and tight detection times of change-points.

stat.ME

Bayesian estimation of large dimensional time varying VARs using copulas

This paper provides a simple, yet reliable, alternative to the (Bayesian) estimation of large multivariate VARs with time variation in the conditional mean equations and/or in the covariance structure. With our new methodology, the original multivariate, n dimensional model is treated as a set of n univariate estimation problems, and cross-dependence is handled through the use of a copula. Thus, only univariate distribution functions are needed when estimating the individual equations, which are often available in closed form, and easy to handle with MCMC (or other techniques). Estimation is carried out in parallel for the individual equations. Thereafter, the individual posteriors are combined with the copula, so obtaining a joint posterior which can be easily resampled. We illustrate our approach by applying it to a large time-varying parameter VAR with 25 macroeconomic variables.

econ.EM