SearcharxivSearch

arXiv subjects

Daniel Salnikov

Publications and source records attributed to Daniel Salnikov.

6 recordsLinked to original sources

Generalized network autoregressive modelling of longitudinal networks with application to presidential elections in the USA

Longitudinal networks are becoming increasingly relevant in the study of dynamic processes characterised by known or inferred community structure. Generalised Network Autoregressive (GNAR) models provide a parsimonious framework for exploiting the underlying network and multivariate time series. We introduce the community-$\alpha$ GNAR model with interactions that exploits prior knowledge or exogenous variables for analysing interactions within and between communities, and can describe serial correlation in longitudinal networks. We derive new explicit finite-sample error bounds that validate analysing high-dimensional longitudinal network data with GNAR models, and provide insights into their attractive properties. We further illustrate our approach by analysing the dynamics of $\textit{Red, Blue}$ and $\textit{Swing}$ states throughout presidential elections in the USA from 1976 to 2020, that is, a time series of length twelve on 51 time series (US states and Washington DC). Our analysis connects network autocorrelation to eight-year long terms, highlights a possible change in the system after the 2016 election, and a difference in behaviour between $\textit{Red}$ and $\textit{Blue}$ states.

stat.ME

The MAPS Algorithm: Fast model-agnostic and distribution-free prediction intervals for supervised learning

A fundamental problem in modern supervised learning is computing reliable conditional prediction intervals in high-dimensional settings: existing methods often rely on restrictive modelling assumptions, do not scale as predictor dimension increases, or only guarantee marginal (population-level) rather than conditional (individual-level) coverage. We introduce the $\textit{lifted predictive model}$ (LPM), a new conditional representation, and propose the MAPS (Model-Agnostic Prediction Sets) algorithm that produces distribution-free conditional prediction intervals and adapts to any trained predictive model. Our procedure is bootstrap-based, scales to high-dimensional inputs and accounts for heteroscedastic errors. We establish the theoretical properties of the LPM, connect prediction accuracy to interval length, and provide sufficient conditions for asymptotic conditional coverage. We evaluate the finite-sample performance of MAPS in a simulation study, and apply our method to simulation-based inference and image classification. In the former, MAPS provides the first approach for debiasing neural Bayes estimators and constructing valid confidence intervals for model parameters given the estimators, at any desired level. In the latter, it provides the first approach that accounts for uncertainty in model calibration and label prediction.

stat.ML

Concentration inequalities for the sample correlation coefficient

The sample correlation coefficient $R$ plays an important role in many statistical analyses. We study the moments of $R$ under the bivariate Gaussian model assumption, provide a novel approximation for its finite sample mean and connect it with known results for the variance. We exploit these approximations to present non-asymptotic concentration inequalities for $R$. Finally, we illustrate our results in a simulation experiment that further validates the approximations presented in this work.

math.ST

Modelling clusters in network time series with an application to presidential elections in the USA

Network time series are becoming increasingly relevant in the study of dynamic processes characterised by a known or inferred underlying network structure. Generalised Network Autoregressive (GNAR) models provide a parsimonious framework for exploiting the underlying network, even in the high-dimensional setting. We extend the GNAR framework by presenting the $\textit{community}$-$\alpha$ GNAR model that exploits prior knowledge and/or exogenous variables for identifying and modelling dynamic interactions across communities in the network. We further analyse the dynamics of $\textit{ Red, Blue}$ and $\textit{Swing}$ states throughout presidential elections in the USA. Our analysis suggests interesting global and communal effects.

stat.ME

New tools for network time series with an application to COVID-19 hospitalisations

Network time series are becoming increasingly important across many areas in science and medicine and are often characterised by a known or inferred underlying network structure, which can be exploited to make sense of dynamic phenomena that are often high-dimensional. For example, the Generalised Network Autoregressive (GNAR) models exploit such structure parsimoniously. We use the GNAR framework to introduce two association measures: the network and partial network autocorrelation functions, and introduce Corbit (correlation-orbit) plots for visualisation. As with regular autocorrelation plots, Corbit plots permit interpretation of underlying correlation structures and, crucially, aid model selection more rapidly than using other tools such as AIC or BIC. We additionally interpret GNAR processes as generalised graphical models, which constrain the processes' autoregressive structure and exhibit interesting theoretical connections to graphical models via utilization of higher-order interactions. We demonstrate how incorporation of prior information is related to performing variable selection and shrinkage in the GNAR context. We illustrate the usefulness of the GNAR formulation, network autocorrelations and Corbit plots by modelling a COVID-19 network time series of the number of admissions to mechanical ventilation beds at 140 NHS Trusts in England & Wales. We introduce the Wagner plot that can analyse correlations over different time periods or with respect to external covariates. In addition, we introduce plots that quantify the relevance and influence of individual nodes. Our modelling provides insight on the underlying dynamics of the COVID-19 series, highlights two groups of geographically co-located `influential' NHS Trusts and demonstrates superior prediction abilities when compared to existing techniques.

stat.ME

A Constructive Proof of the Glivenko-Cantelli Theorem

The Glivenko-Cantelli theorem states that the empirical distribution function converges uniformly almost surely to the theoretical distribution for a random variable $X \in \mathbb{R}$. This is an important result because it establishes the fact that sampling does capture the dispersion measure the distribution function $F$ imposes. In essence, sampling permits one to learn and infer the behavior of $F$ by only looking at observations from $X$. The probabilities that are inferred from samples $\mathbf{X}$ will become more precise as the sample size increases and more data becomes available. Therefore, it is valid to study distributions via samples. The proof present here is constructive, meaning that the result is derived directly from the fact that the empirical distribution function converges pointwise almost surely to the theoretical distribution. The work includes a proof of this preliminary statement and attempts to motivate the intuition one gets from sampling techniques when studying the regions in which a model concentrates probability. The sets where dispersion is described with precision by the empirical distribution function will eventually cover the entire sample space.

math.PR