SearcharxivSearch

arXiv subjects

Matteo Barigozzi

Publications and source records attributed to Matteo Barigozzi.

At least 37 records · Page 2Linked to original sources

Dynamic Factor Models: a Genealogy

Dynamic factor models have been developed out of the need of analyzing and forecasting time series in increasingly high dimensions. While mathematical statisticians faced with inference problems in high-dimensional observation spaces were focusing on the so-called spiked-model-asymptotics, econometricians adopted an entirely and considerably more effective asymptotic approach, rooted in the factor models originally considered in psychometrics. The so-called dynamic factor model methods, in two decades, has grown into a wide and successful body of techniques that are widely used in central banks, financial institutions, economic and statistical institutes. The objective of this chapter is not an extensive survey of the topic but a sketch of its historical growth, with emphasis on the various assumptions and interpretations, and a family tree of its main variants.

econ.EM

General Spatio-Temporal Factor Models for High-Dimensional Random Fields on a Lattice

Motivated by the need for analysing large spatio-temporal panel data, we introduce a novel dimensionality reduction methodology for $n$-dimensional random fields observed across a number $S$ spatial locations and $T$ time periods. We call it General Spatio-Temporal Factor Model (GSTFM). First, we provide the probabilistic and mathematical underpinning needed for the representation of a random field as the sum of two components: the common component (driven by a small number $q$ of latent factors) and the idiosyncratic component (mildly cross-correlated). We show that the two components are identified as $n\to\infty$. Second, we propose an estimator of the common component and derive its statistical guarantees (consistency and rate of convergence) as $\min(n, S, T )\to\infty$. Third, we propose an information criterion to determine the number of factors. Estimation makes use of Fourier analysis in the frequency domain and thus we fully exploit the information on the spatio-temporal covariance structure of the whole panel. Synthetic data examples illustrate the applicability of GSTFM and its advantages over the extant generalized dynamic factor model that ignores the spatial correlations.

stat.ME

Robust Tensor Factor Analysis

We consider (robust) inference in the context of a factor model for tensor-valued sequences. We study the consistency of the estimated common factors and loadings space when using estimators based on minimising quadratic loss functions. Building on the observation that such loss functions are adequate only if sufficiently many moments exist, we extend our results to the case of heavy-tailed distributions by considering estimators based on minimising the Huber loss function, which uses an $L_{1}$ -norm weight on outliers. We show that such class of estimators is robust to the presence of heavy tails, even when only the second moment of the data exists. We also propose a modified version of the eigenvalue-ratio principle to estimate the dimensions of the core tensor and show the consistency of the resultant estimators without any condition on the relative rates of divergence of the sample size and dimensions. Extensive numerical studies are conducted to show the advantages of the proposed methods over the state-of-the-art ones especially under the heavy-tailed cases. An import/export dataset of a variety of commodities across multiple countries is analyzed to show the practical usefulness of the proposed robust estimation procedure. An R package ``RTFA" implementing the proposed methods is available on R CRAN.

stat.ME

fnets: An R Package for Network Estimation and Forecasting via Factor-Adjusted VAR Modelling

The package fnets for the R language implements the suite of methodologies proposed by Barigozzi et al. (2022) for the network estimation and forecasting of high-dimensional time series under a factor-adjusted vector autoregressive model, which permits strong spatial and temporal correlations in the data. Additionally, we provide tools for visualising the networks underlying the time series data after adjusting for the presence of factors. The package also offers data-driven methods for selecting tuning parameters including the number of factors, vector autoregressive order and thresholds for estimating the edge sets of the networks of interest in time series analysis. We demonstrate various features of fnets on simulated datasets as well as real data on electricity prices.

stat.CO

Multidimensional dynamic factor models

This paper generalises dynamic factor models for multidimensional dependent data. In doing so, it develops an interpretable technique to study complex information sources ranging from repeated surveys with a varying number of respondents to panels of satellite images. We specialise our results to model microeconomic data on US households jointly with macroeconomic aggregates. This results in a powerful tool able to generate localised predictions, counterfactuals and impulse response functions for individual households, accounting for traditional time-series complexities depicted in the state-space literature. The model is also compatible with the growing focus of policymakers for real-time economic analysis as it is able to process observations online, while handling missing values and asynchronous data releases.

econ.EM

Inference in heavy-tailed non-stationary multivariate time series

We study inference on the common stochastic trends in a non-stationary, $N$-variate time series $y_{t}$, in the possible presence of heavy tails. We propose a novel methodology which does not require any knowledge or estimation of the tail index, or even knowledge as to whether certain moments (such as the variance) exist or not, and develop an estimator of the number of stochastic trends $m$ based on the eigenvalues of the sample second moment matrix of $y_{t}$. We study the rates of such eigenvalues, showing that the first $m$ ones diverge, as the sample size $T$ passes to infinity, at a rate faster by $O\left(T \right)$ than the remaining $N-m$ ones, irrespective of the tail index. We thus exploit this eigen-gap by constructing, for each eigenvalue, a test statistic which diverges to positive infinity or drifts to zero according to whether the relevant eigenvalue belongs to the set of the first $m$ eigenvalues or not. We then construct a randomised statistic based on this, using it as part of a sequential testing procedure, ensuring consistency of the resulting estimator of $m$. We also discuss an estimator of the common trends based on principal components and show that, up to a an invertible linear transformation, such estimator is consistent in the sense that the estimation error is of smaller order than the trend itself. Finally, we also consider the case in which we relax the standard assumption of \textit{i.i.d.} innovations, by allowing for heterogeneity of a very general form in the scale of the innovations. A Monte Carlo study shows that the proposed estimator for $m$ performs particularly well, even in samples of small size. We complete the paper by presenting four illustrative applications covering commodity prices, interest rates data, long run PPP and cryptocurrency markets.

econ.EM

An algebraic estimator for large spectral density matrices

We propose a new estimator of high-dimensional spectral density matrices, called UNshrunk ALgebraic Spectral Estimator (UNALSE), under the assumption of an underlying low rank plus sparse structure, as typically assumed in dynamic factor models. The UNALSE is computed by minimizing a quadratic loss under a nuclear norm plus $l_1$ norm constraint to control the latent rank and the residual sparsity pattern. The loss function requires as input the classical smoothed periodogram estimator and two threshold parameters, the choice of which is thoroughly discussed. We prove consistency of UNALSE as both the dimension $p$ and the sample size $T$ diverge to infinity, as well as algebraic consistency, i.e., the recovery of latent rank and residual sparsity pattern with probability one. The finite sample properties of UNALSE are studied by means of an extended simulation exercise as well as an empirical analysis of US macroeconomic data.

math.ST

Consistent estimation of high-dimensional factor models when the factor number is over-estimated

A high-dimensional $r$-factor model for an $n$-dimensional vector time series is characterised by the presence of a large eigengap (increasing with $n$) between the $r$-th and the $(r+1)$-th largest eigenvalues of the covariance matrix. Consequently, Principal Component (PC) analysis is the most popular estimation method for factor models and its consistency, when $r$ is correctly estimated, is well-established in the literature. However, popular factor number estimators often suffer from the lack of an obvious eigengap in empirical eigenvalues and tend to over-estimate $r$ due, for example, to the existence of non-pervasive factors affecting only a subset of the series. We show that the errors in the PC estimators resulting from the over-estimation of $r$ are non-negligible, which in turn lead to the violation of the conditions required for factor-based large covariance estimation. To remedy this, we propose new estimators of the factor model based on scaling the entries of the sample eigenvectors. We show both theoretically and numerically that the proposed estimators successfully control for the over-estimation error, and investigate their performance when applied to risk minimisation of a portfolio of financial time series.

stat.ME

Large-Dimensional Dynamic Factor Models: Estimation of Impulse-Response Functions with $I(1)$ Cointegrated Factors

We study a large-dimensional Dynamic Factor Model where: (i)~the vector of factors $\mathbf F_t$ is $I(1)$ and driven by a number of shocks that is smaller than the dimension of $\mathbf F_t$; and, (ii)~the idiosyncratic components are either $I(1)$ or $I(0)$. Under~(i), the factors $\mathbf F_t$ are cointegrated and can be modeled as a Vector Error Correction Model (VECM). Under (i) and (ii), we provide consistent estimators, as both the cross-sectional size $n$ and the time dimension $T$ go to infinity, for the factors, the loadings, the shocks, the coefficients of the VECM and therefore the Impulse-Response Functions (IRF) of the observed variables to the shocks.~Furthermore: possible deterministic linear trends are fully accounted for, and the case of an unrestricted VAR in the levels $\mathbf F_t$, instead of a VECM, is also studied. The finite-sample properties the proposed estimators are explored by means of a MonteCarlo exercise. Finally, we revisit two distinct and widely studied empirical applications. By correctly modeling the long-run dynamics of the factors, our results partly overturn those obtained by recent literature. Specifically, we find that: (i) oil price shocks have just a temporary effect on US real activity; and, (ii) in response to a positive news shock, the economy first experiences a significant boom, and then a milder recession.

stat.ME

Sequential testing for structural stability in approximate factor models

We develop a monitoring procedure to detect changes in a large approximate factor model. Letting $r$ be the number of common factors, we base our statistics on the fact that the $\left( r+1\right) $-th eigenvalue of the sample covariance matrix is bounded under the null of no change, whereas it becomes spiked under changes. Given that sample eigenvalues cannot be estimated consistently under the null, we randomise the test statistic, obtaining a sequence of \textit{i.i.d} statistics, which are used for the monitoring scheme. Numerical evidence shows a very small probability of false detections, and tight detection times of change-points.

stat.ME

Quasi Maximum Likelihood Estimation of Non-Stationary Large Approximate Dynamic Factor Models

This paper considers estimation of large dynamic factor models with common and idiosyncratic trends by means of the Expectation Maximization algorithm, implemented jointly with the Kalman smoother. We show that, as the cross-sectional dimension $n$ and the sample size $T$ diverge to infinity, the common component for a given unit estimated at a given point in time is $\min(\sqrt n,\sqrt T)$-consistent. The case of local levels and/or local linear trends trends is also considered. By means of a MonteCarlo simulation exercise, we compare our approach with estimators based on principal component analysis.

econ.EM

Generalized Dynamic Factor Models and Volatilities: Consistency, rates, and prediction intervals

Volatilities, in high-dimensional panels of economic time series with a dynamic factor structure on the levels or returns, typically also admit a dynamic factor decomposition. We consider a two-stage dynamic factor model method recovering the common and idiosyncratic components of both levels and log-volatilities. Specifically, in a first estimation step, we extract the common and idiosyncratic shocks for the levels, from which a log-volatility proxy is computed. In a second step, we estimate a dynamic factor model, which is equivalent to a multiplicative factor structure for volatilities, for the log-volatility panel. By exploiting this two-stage factor approach, we build one-step-ahead conditional prediction intervals for large $n \times T$ panels of returns. Those intervals are based on empirical quantiles, not on conditional variances; they can be either equal- or unequal- tailed. We provide uniform consistency and consistency rates results for the proposed estimators as both $n$ and $T$ tend to infinity. We study the finite-sample properties of our estimators by means of Monte Carlo simulations. Finally, we apply our methodology to a panel of asset returns belonging to the S&P100 index in order to compute one-step-ahead conditional prediction intervals for the period 2006-2013. A comparison with the componentwise GARCH benchmark (which does not take advantage of cross-sectional information) demonstrates the superiority of our approach, which is genuinely multivariate (and high-dimensional), nonparametric, and model-free.

econ.EM

Maximum entropy approaches for the study of triadic motifs in the Mergers & Acquisitions network

In the past years statistical physics has been successfully applied for complex networks modelling. In particular, it has been shown that the maximum entropy principle can be exploited in order to construct graph ensembles for real-world networks which maximize the randomness of the graph structure keeping fixed some topological constraint. Such ensembles can be used as null models to detect statistically significant structural patterns and to reconstruct the network structure in cases of incomplete information. Recently, these randomizing methods have been used for the study of self-organizing systems in economics and finance, such as interbank and world trade networks, in order to detect topological changes and, possibly, early-warning signals for the economical crisis. In this work we consider the configuration models with different constraints for the network of mergers and acquisitions (M&As), Comparing triadic and dyadic motifs, for both the binary and weighted M&A network, with the randomized counterparts can shed light on its organization at higher order level.

physics.soc-ph

Determining the dimension of factor structures in non-stationary large datasets

We propose a procedure to determine the dimension of the common factor space in a large, possibly non-stationary, dataset. Our procedure is designed to determine whether there are (and how many) common factors (i) with linear trends, (ii) with stochastic trends, (iii) with no trends, i.e. stationary. Our analysis is based on the fact that the largest eigenvalues of a suitably scaled covariance matrix of the data (corresponding to the common factor part) diverge, as the dimension $N$ of the dataset diverges, whilst the others stay bounded. Therefore, we propose a class of randomised test statistics for the null that the $p$-th eigenvalue diverges, based directly on the estimated eigenvalue. The tests only requires minimal assumptions on the data, and no restrictions on the relative rates of divergence of $N$ and $T$ are imposed. Monte Carlo evidence shows that our procedure has very good finite sample properties, clearly dominating competing approaches when no common factors are present. We illustrate our methodology through an application to US bond yields with different maturities observed over the last 30 years. A common linear trend and two common stochastic trends are found and identified as the classical level, slope and curvature factors.

stat.ME

Simultaneous multiple change-point and factor analysis for high-dimensional time series

We propose the first comprehensive treatment of high-dimensional time series factor models with multiple change-points in their second-order structure. We operate under the most flexible definition of piecewise stationarity, and estimate the number and locations of change-points consistently as well as identifying whether they originate in the common or idiosyncratic components. Through the use of wavelets, we transform the problem of change-point detection in the second-order structure of a high-dimensional time series, into the (relatively easier) problem of change-point detection in the means of high-dimensional panel data. Also, our methodology circumvents the difficult issue of the accurate estimation of the true number of factors in the presence of multiple change-points by adopting a screening procedure. We further show that consistent factor analysis is achieved over each segment defined by the change-points estimated by the proposed methodology. In extensive simulation studies, we observe that factor analysis prior to change-point detection improves the detectability of change-points, and identify and describe an interesting `spillover' effect in which substantial breaks in the idiosyncratic components get, naturally enough, identified as change-points in the common components, which prompts us to regard the corresponding change-points as also acting as a form of `factors'. Our methodology is implemented in the R package {\tt factorcpt}, available from CRAN.

stat.ME

Common factors, trends, and cycles in large datasets

This paper considers a non-stationary dynamic factor model for large datasets to disentangle long-run from short-run co-movements. We first propose a new Quasi Maximum Likelihood estimator of the model based on the Kalman Smoother and the Expectation Maximisation algorithm. The asymptotic properties of the estimator are discussed. Then, we show how to separate trends and cycles in the factors by mean of eigenanalysis of the estimated non-stationary factors. Finally, we employ our methodology on a panel of US quarterly macroeconomic indicators to estimate aggregate real output, or Gross Domestic Output, and the output gap.

stat.ME

Spatio-Temporal Patterns of the International Merger and Acquisition Network

This paper analyses the world web of mergers and acquisitions (M&As) using a complex network approach. We use data of M&As to build a temporal sequence of binary and weighted-directed networks for the period 1995-2010 and 224 countries (nodes) connected according to their M&As flows (links). We study different geographical and temporal aspects of the international M&A network (IMAN), building sequences of filtered sub-networks whose links belong to specific intervals of distance or time. Given that M&As and trade are complementary ways of reaching foreign markets, we perform our analysis using statistics employed for the study of the international trade network (ITN), highlighting the similarities and differences between the ITN and the IMAN. In contrast to the ITN, the IMAN is a low density network characterized by a persistent giant component with many external nodes and low reciprocity. Clustering patterns are very heterogeneous and dynamic. High-income economies are the main acquirers and are characterized by high connectivity, implying that most countries are targets of a few acquirers. Like in the ITN, geographical distance strongly impacts the structure of the IMAN: link-weights and node degrees have a non-linear relation with distance, and an assortative pattern is present at short distances.

physics.soc-ph

Dynamic Factor Models, Cointegration, and Error Correction Mechanisms

The paper studies Non-Stationary Dynamic Factor Models such that the factors $\mathbf F_t$ are $I(1)$ and singular, i.e. $\mathbf F_t$ has dimension $r$ and is driven by a $q$-dimensional white noise, the common shocks, with $q<r$. We show that $\mathbf F_t$ is driven by $r-c$ permanent shocks, where $c$ is the cointegration rank of $\mathbf F_t$, and $q-(r-c)<c$ transitory shocks, thus the same result as in the non-singular case for the permanent shocks but not for the transitory shocks. Our main result is obtained by combining the classic Granger Representation Theorem with recent results by Anderson and Deistler on singular stochastic vectors: if $(1-L)\mathbf F_t$ is singular and has {\it rational} spectral density then, for generic values of the parameters, $\mathbf F_t$ has an autoregressive representation with a {\it finite-degree} matrix polynomial fulfilling the restrictions of a Vector Error Correction Mechanism with $c$ error terms. This result is the basis for consistent estimation of Non-Stationary Dynamic Factor Models. The relationship between cointegration of the factors and cointegration of the observable variables is also discussed.

math.ST