Searcharxiv⌕ Search

arXiv subjects

Davide La Vecchia

Publications and source records attributed to Davide La Vecchia.

12 recordsLinked to original sources

Flexible latent variable models on graphs: Laplace approximated inference for multiview network data

We propose a novel and flexible nonlinear approach for dimensionality reduction of large-scale multiview network data and derive its theory. The (linear) predictor incorporates observed covariates (edge-specific, layer-specific, and global) and Gaussian latent factors. Inference is conducted via the graph Laplace approximated maximum likelihood estimator. Letting $K$ denote the number of network layers and $n_V$ the number of nodes, we derive asymptotic theory under two regimes: (i) $K \to \infty$ with fixed $n_V$, and (ii) double asymptotics $K, n_V \to \infty$, establishing consistency and asymptotic normality for local and global parameters, with distinct convergence rates. In an application to the gravity model for commodity trades, we use a zero-adjusted Gamma distribution with latent factors and observable covariates (e.g.\ distance, tariffs, common language) to capture excess zeros, skewness, and unobserved heterogeneity. Synthetic and real-data exercises show that our approach outperforms the routinely applied Poisson pseudo-maximum likelihood estimator with fixed effects. We complement our theoretical and empirical contributions with open-source R/C++ routines and a novel strategy for starting values selection.

stat.ME↗

E-ROBOT: a dimension-free method for robust statistics and machine learning via Schrödinger bridge

We propose the Entropic-regularized Robust Optimal Transport (E-ROBOT) framework, a novel method that combines the robustness of ROBOT with the computational and statistical benefits of entropic regularization. We show that, rooted in the Schrödinger bridge problem theory, E-ROBOT defines the robust Sinkhorn divergence $\overline{W}_{\varepsilon,λ}$, where the parameter $λ$ controls robustness and $\varepsilon$ governs the regularization strength. Letting $n\in \mathbb{N}$ denote the sample size, a central theoretical contribution is establishing that the sample complexity of $\overline{W}_{\varepsilon,λ}$ is $\mathcal{O}(n^{-1/2})$, thereby avoiding the curse of dimensionality that plagues standard ROBOT. This dimension-free property unlocks the use of $\overline{W}_{\varepsilon,λ}$ as a loss function in large-dimensional statistical and machine learning tasks. With this regard, we demonstrate its utility through four applications: goodness-of-fit testing; computation of barycenters for corrupted 2D and 3D shapes; definition of gradient flows; and image colour transfer. From the computation standpoint, a perk of our novel method is that it can be easily implemented by modifying existing (\texttt{Python}) routines. From the theoretical standpoint, our work opens the door to many research directions in statistics and machine learning: we discuss some of them.

stat.ML↗

Discussion of "Robust Distance Covariance" by S. Leyder, J. Raymaekers, and P.J. Rousseeuw

Distance covariance and distance correlation have long been regarded as natural measures of dependence between two random vectors, and have been used in a variety of situations for testing independence. Despite their popularity, the robustness of their empirical versions remain highly undiscovered. The paper named "Robust Distance Covariance" by S. Leyder, J. Raymaekers, and P.J. Rousseeuw (below referred to as [LRR]), which this article is discussing about, has provided a welcome addition to the literature. Among some intriguing results in [LRR], we find ourselves particularly interested in the so-called "robustness by transformation" that was highlighted when they used a clever trick named "the biloop transformation" to obtain a bounded and redescending influence function. Building on the measure-transportation-based notions of directional ranks and signs, we show how the "robustness via transformation" principle emphasized by [LRR] extends beyond the case of bivariate independence that [LRR] has investigated and also applies in higher-dimension Euclidean spaces and on compact manifolds. The case of directional variables (taking values on (hyper)spheres) is given special attention.

stat.ME↗

On the use of the cumulant generating function for inference on time series

We introduce innovative inference procedures for analyzing time series data. Our methodology enables density approximation and composite hypothesis testing based on Whittle's estimator, a widely applied M-estimator in the frequency domain. Its core feature involves the (general Legendre transform of the) cumulant generating function of the Whittle likelihood score, as obtained using an approximated distribution of the periodogram ordinates. We present a testing algorithm that significantly expands the applicability of the state-of-the-art saddlepoint test, while maintaining the numerical accuracy of the saddlepoint approximation. Additionally, we demonstrate connections between our findings and three other prevalent frequency domain approaches: the bootstrap, empirical likelihood, and exponential tilting. Numerical examples using both simulated and real data illustrate the advantages and accuracy of our methodology.

stat.ME↗

Inference via robust optimal transportation: theory and methods

Optimal transportation theory and the related $p$-Wasserstein distance ($W_p$, $p\geq 1$) are widely-applied in statistics and machine learning. In spite of their popularity, inference based on these tools has some issues. For instance, it is sensitive to outliers and it may not be even defined when the underlying model has infinite moments. To cope with these problems, first we consider a robust version of the primal transportation problem and show that it defines the {robust Wasserstein distance}, $W^{(λ)}$, depending on a tuning parameter $λ> 0$. Second, we illustrate the link between $W_1$ and $W^{(λ)}$ and study its key measure theoretic aspects. Third, we derive some concentration inequalities for $W^{(λ)}$. Fourth, we use $W^{(λ)}$ to define minimum distance estimators, we provide their statistical guarantees and we illustrate how to apply the derived concentration inequalities for a data driven selection of $λ$. Fifth, we provide the {dual} form of the robust optimal transportation problem and we apply it to machine learning problems (generative adversarial networks and domain adaptation). Numerical exercises provide evidence of the benefits yielded by our novel methods.

math.ST↗

General Spatio-Temporal Factor Models for High-Dimensional Random Fields on a Lattice

Motivated by the need for analysing large spatio-temporal panel data, we introduce a novel dimensionality reduction methodology for $n$-dimensional random fields observed across a number $S$ spatial locations and $T$ time periods. We call it General Spatio-Temporal Factor Model (GSTFM). First, we provide the probabilistic and mathematical underpinning needed for the representation of a random field as the sum of two components: the common component (driven by a small number $q$ of latent factors) and the idiosyncratic component (mildly cross-correlated). We show that the two components are identified as $n\to\infty$. Second, we propose an estimator of the common component and derive its statistical guarantees (consistency and rate of convergence) as $\min(n, S, T )\to\infty$. Third, we propose an information criterion to determine the number of factors. Estimation makes use of Fourier analysis in the frequency domain and thus we fully exploit the information on the spatio-temporal covariance structure of the whole panel. Synthetic data examples illustrate the applicability of GSTFM and its advantages over the extant generalized dynamic factor model that ignores the spatial correlations.

stat.ME↗

Some novel aspects of quantile regression: local stationarity, random forests and optimal transportation

This paper is written for a Festschrift in honour of Professor Marc Hallin and it proposes some developments on quantile regression. We connect our investigation to Marc's scientific production and we present some theoretical and methodological advances for quantiles estimation in non standard settings. We split our contributions in two parts. The first part is about conditional quantiles estimation for nonstationary time series. The second part is about conditional quantiles estimation for the analysis of multivariate independent data in the presence of possibly large dimensional covariates. Monte Carlo studies illustrate numerically the performance of our methods and compare them to some extant techniques.

stat.AP↗

GLAMLE: inference for multiview network data in the presence of latent variables, with application to commodities trading

The statistical analysis of import/export data is helpful to understand the mechanism that determines exchanges in an economic network. The probability of having a commercial relationship between two countries often depends on some unobservable (or not easy-to-measure) factors, like socio-economical conditions, political views, level of the infrastructures. To conduct inference on this type of data, we introduce a novel class of latent variable models for multiview networks, where a multivariate latent Gaussian variable determines the probabilistic behavior of the edges. We label our model the Graph Generalized Linear Latent Variable Model (GGLLVM) and we base our inference on the maximization of the Laplace-approximated likelihood. We call the resulting M-estimator the Graph Laplace-Approximated Maximum Likelihood Estimator (GLAMLE) and we study its statistical properties. Numerical experiments on simulated networks illustrate that the GLAMLE yields fast and accurate inference. A real data application to commodities trading in Central Europe countries unveils the import/export propensity that each node of the network has toward other nodes, along with additional information specific to each traded commodity.

stat.ME↗

A Higher-Order Correct Fast Moving-Average Bootstrap for Dependent Data

We develop and implement a novel fast bootstrap for dependent data. Our scheme is based on the i.i.d. resampling of the smoothed moment indicators. We characterize the class of parametric and semi-parametric estimation problems for which the method is valid. We show the asymptotic refinements of the proposed procedure, proving that it is higher-order correct under mild assumptions on the time series, the estimating functions, and the smoothing kernel. We illustrate the applicability and the advantages of our procedure for Generalized Empirical Likelihood estimation. As a by-product, our fast bootstrap provides higher-order correct asymptotic confidence distributions. Monte Carlo simulations on an autoregressive conditional duration model provide numerical evidence that the novel bootstrap yields higher-order accurate confidence intervals. A real-data application on dynamics of trading volume of stocks illustrates the advantage of our method over the routinely-applied first-order asymptotic theory, when the underlying distribution of the test statistic is skewed or fat-tailed.

stat.ME↗

Saddlepoint approximations for spatial panel data models

We develop new higher-order asymptotic techniques for the Gaussian maximum likelihood estimator in a spatial panel data model, with fixed effects, time-varying covariates, and spatially correlated errors. Our saddlepoint density and tail area approximation feature relative error of order $O(1/(n(T-1)))$ with $n$ being the cross-sectional dimension and $T$ the time-series dimension. The main theoretical tool is the tilted-Edgeworth technique in a non-identically distributed setting. The density approximation is always non-negative, does not need resampling, and is accurate in the tails. Monte Carlo experiments on density approximation and testing in the presence of nuisance parameters illustrate the good performance of our approximation over first-order asymptotics and Edgeworth expansions. An empirical application to the investment-saving relationship in OECD (Organisation for Economic Co-operation and Development) countries shows disagreement between testing results based on first-order asymptotics and saddlepoint techniques.

math.ST↗

Rank-Based Testing for Semiparametric VAR Models: a measure transportation approach

We develop a class of tests for semiparametric vector autoregressive (VAR) models with unspecified innovation densities, based on the recent measure-transportation-based concepts of multivariate {\it center-outward ranks} and {\it signs}. We show that these concepts, combined with Le Cam's asymptotic theory of statistical experiments, yield novel testing procedures, which (a)~are valid under a broad class of innovation densities (possibly non-elliptical, skewed, and/or with infinite moments), (b)~are optimal (locally asymptotically maximin or most stringent) at selected ones, and (c) are robust against additive outliers. In order to do so, we establish a H\' ajek asymptotic representation result, of independent interest, for a general class of center-outward rank-based serial statistics. As an illustration, we consider the problems of testing the absence of serial correlation in multiple-output and possibly non-linear regression (an extension of the classical Durbin-Watson problem) and the sequential identification of the order $p$ of a vector autoregressive (VAR($p$)) model. A Monte Carlo comparative study of our tests and their routinely-applied Gaussian competitors demonstrates the benefits (in terms of size, power, and robustness) of our methodology; these benefits are particularly significant in the presence of asymmetric and leptokurtic innovation densities. A real data application concludes the paper.

math.ST↗

Center-Outward R-Estimation for Semiparametric VARMA Models

We propose a new class of R-estimators for semiparametric VARMA models in which the innovation density plays the role of the nuisance parameter. Our estimators are based on the novel concepts of multivariate center-outward ranks and signs. We show that these concepts, combined with Le Cam's asymptotic theory of statistical experiments, yield a class of semiparametric estimation procedures, which are efficient (at a given reference density), root-$n$ consistent, and asymptotically normal under a broad class of (possibly non elliptical) actual innovation densities. No kernel density estimation is required to implement our procedures. A Monte Carlo comparative study of our R-estimators and other routinely-applied competitors demonstrates the benefits of the novel methodology, in large and small sample. Proofs, computational aspects, and further numerical results are available in the supplementary material.

math.ST↗