SearcharxivSearch

arXiv subjects

Daria Tieplova

Publications and source records attributed to Daria Tieplova.

7 recordsLinked to original sources

Change Point Detection and Localization in High-Dimensional Time Series

We present new inference tools for change point detection in high-dimensional time series. We discuss two distinct statistical applications: First, sequential change point testing in an incoming data-stream. Second, retrospective localization of multiple changes, with confidence intervals at a globally controlled error level. Test statistics are built on the maximum norm to generate power against sparse and asynchronous changes. Both problems are tackled by related multiscale statistics that search for changes in the data at many different levels of resolution. For fixed dimension, our statistical approaches can be validated using traditional Hölderian invariance principles. In this paper, we present the high-dimensional analogue: Hölder-Gauss-approximations, which can be (roughly) interpreted as the Gaussian approximation for a Hölder-norm of the high-dimensional partial sum process. Such approximations are of interest beyond change point detection and can be used for other problems such as for stationarity testing in high dimensions. We evaluate finite-sample performance in a simulation study and give an application to air contamination due to wildfires in California, which occurs asynchronously across a panel of measuring stations.

stat.ME

Information-theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime

We rigorously analyse fully-trained neural networks of arbitrary depth in the Bayesian optimal setting in the so-called proportional scaling regime where the number of training samples and width of the input and all inner layers diverge proportionally. We prove an information-theoretic equivalence between the Bayesian deep neural network model trained from data generated by a teacher with matching architecture, and a simpler model of optimal inference in a generalized linear model. This equivalence enables us to compute the optimal generalization error for deep neural networks in this regime. We thus prove the "deep Gaussian equivalence principle" conjectured in Cui et al. (2023) (arXiv:2302.00375). Our result highlights that in order to escape this "trivialisation" of deep neural networks (in the sense of reduction to a linear model) happening in the strongly overparametrized proportional regime, models trained from much more data have to be considered.

math.ST

Information-theoretic limits and approximate message-passing for high-dimensional time series

High-dimensional time series appear in many scientific setups, demanding a nuanced approach to model and analyze the underlying dependence structure. Theoretical advancements so far often rely on stringent assumptions regarding the sparsity of the underlying signal. In non-sparse regimes, analyses have primarily focused on linear regression models with the design matrix having independent rows. In this paper, we expand the scope by investigating a high-dimensional time series model wherein the number of features grows proportionally to the number of sampling points, without assuming sparsity in the signal. Specifically, we consider the stochastic regression model and derive a single-letter formula for the normalized mutual information between observations and the signal, as well as for minimum mean-square errors. We also empirically study the vector approximate message passing VAMP algorithm and show that, despite the lack of theoretical guarantees, its performance for inference in our time series model is robust and often statistically optimal.

cs.IT

Fundamental limits of overparametrized shallow neural networks for supervised learning

We carry out an information-theoretical analysis of a two-layer neural network trained from input-output pairs generated by a teacher network with matching architecture, in overparametrized regimes. Our results come in the form of bounds relating i) the mutual information between training data and network weights, or ii) the Bayes-optimal generalization error, to the same quantities but for a simpler (generalized) linear model for which explicit expressions are rigorously known. Our bounds, which are expressed in terms of the number of training samples, input dimension and number of hidden units, thus yield fundamental performance limits for any neural network (and actually any learning procedure) trained from limited data generated according to our two-layer teacher neural network model. The proof relies on rigorous tools from spin glasses and is guided by ``Gaussian equivalence principles'' lying at the core of numerous recent analyses of neural networks. With respect to the existing literature, which is either non-rigorous or restricted to the case of the learning of the readout weights only, our results are information-theoretic (i.e. are not specific to any learning algorithm) and, importantly, cover a setting where all the network parameters are trained.

cs.LG

On the largest singular values of certain large random matrices with application to the estimation of the minimal dimension of the state-space representations of high-dimensional time series

This paper is devoted to the estimation of the minimal dimension P of the state-space realizations of a high-dimensional time series y, defined as a noisy version (the noise is white and Gaussian) of a useful signal with low rank rational spectral density, in the high-dimensional asymptotic regime where the number of available samples N and the dimension of the time series M converge towards infinity at the same rate. In the classical low-dimensional regime, P is estimated as the number of significant singular values of the empirical autocovariance matrix between the past and the future of y, or as the number of significant estimated canonical correlation coefficients between the past and the future of y. Generalizing large random matrix methods developed in the past to analyze classical spiked models, the behaviour of the above singular values and canonical correlation coefficients is studied in the high-dimensional regime. It is proved that they are smaller than certain thresholds depending on the statistics of the noise, except a finite number of outliers that are due to the useful signal. The number of singular values of the sample autocovariance matrix above the threshold is evaluated, is shown to be almost independent from P in general, and cannot therefore be used to estimate P accurately. In contrast, the number s of canonical correlation coefficients larger than the corresponding threshold is shown to be less than or equal to P, and explicit conditions under which it is equal to P are provided. Under the corresponding assumptions, s is thus a consistent estimate of P in the high-dimensional regime. The core of the paper is the development of the necessary large random matrix tools.

cs.IT

On the behaviour of large empirical autocovariance matrices between the past and the future

The asymptotic behaviour of the distribution of the squared singular values of the sample autocovariance matrix between the past and the future of a high-dimensional complex Gaussian uncorrelated sequence is studied. Using Gaussian tools, it is established the distribution behaves as a deterministic probability measure whose support S is characterized. It is also established that the singular values to the square are almost surely located in a neighbourhood of S.

math.PR

Distribution of eigenvalues of sample covariance matrices with tensor product samples

We consider $n^2\times n^2$ real symmetric and hermitian matrices $M_n$, which are equal to sum of $m_n$ tensor products of vectors $X^μ=B(Y^μ\otimes Y^μ)$, $μ=1,\dots,m_n$, where $Y^μ$ are i.i.d. random vectors from $\mathbb R^n (\mathbb C^n)$ with zero mean and unit variance of components, and $B$ is an $n^2\times n^2$ positive definite non-random matrix. We prove that if $m_n/n^2\to c\in [0,+\infty)$ and the Normalized Counting Measure of eigenvalues of $BJB$, where $J$ is defined below, converges weakly, then the Normalized Counting Measure of eigenvalues of $M_n$ converges weakly in probability to a non-random limit and its Stieltjes transform can be found from a certain functional equation.

math-ph