SearcharxivSearch

arXiv subjects

Veronica Vinciotti

Publications and source records attributed to Veronica Vinciotti.

At least 19 recordsLinked to original sources

Gaussian Graphical Models for Partially Observed Multivariate Functional Data

In many applications, the variables that characterize a stochastic system are measured along a second dimension, such as time. This results in multivariate functional data and the interest is in describing the statistical dependencies among these variables. It is often the case that the functional data are only partially observed. This creates additional challenges to statistical inference, since the functional principal component scores, which capture all the information from these data, cannot be computed. Under an assumption of Gaussianity and of partial separability of the covariance operator, we develop an Expectation-Maximization (EM)-type algorithm for penalized inference of a functional graphical model from multivariate functional data which are only partially observed. A simulation study and an illustration on environmental, social and governance (ESG) data show the potential of the proposed method.

stat.ME

Loglinear modelling of huge contingency tables

Contingency tables are the canonical representation of multivariate categorical data. As the size of the contingency table grows exponentially with the number of variables, even a moderate number of variables, each with a moderate number of levels, results in a huge number of cells, the majority of which remains empty even with a significant amount of data. We propose efficient methods for inferring higher-order loglinear models by performing subsampling on the set of the empty cells. First, we derive the likelihood under a zero-deflated Poisson sampling scheme. This is maximized via an efficient iteratively re-weighted least squares algorithm, leading to consistent and close to efficient estimators. This method works well for moderately sized contingency tables, but runs into computational instability when the number of dimensions grows. By sacrificing some efficiency, we show that nested case-control multinomial sampling combined with a degenerate logistic regression approach is also consistent and can be applied to arbitrarily large contingency tables. We illustrate the method with an analysis of data from the General Social Survey, which consists of $15014$ observations in a $69$-dimensional contingency table with a total of $6.6\times 10^{38}$ cells.

stat.ME

Causal invariance in graphical models with latent variables

Causal discovery aims to identify causal relationships among variables from observational or interventional data, typically represented by a directed acyclic graph (DAG). The causal invariance principle enables the identification of the causal parents of target variables by exploiting the stability of causal effects across different experimental settings. When some parents are unobserved, however, the induced graph over the observed variables may no longer be a DAG, and it may not be unique, complicating causal inference. For relevant configurations of latent parents, we characterize the induced graph and formalize the conditions under which causal invariance is preserved for the identification of the observed parents. Necessary and sufficient conditions for testing such invariance are formally established for a multivariate Gaussian target.

stat.ME

Bayesian nonparametric Mallows model for clustering preference data

Preference learning refers to the learning of latent patterns from ranking and preference data of different kinds. Typical aims of preference learning are to infer a shared consensus ranking, to learn individual-level preferences, and to perform unsupervised clustering. The Mallows model is among the few approaches that can achieve all these objectives jointly. Previous work has developed computationally tractable methods for Bayesian inference based on a MCMC Metropolis-Hastings scheme, where clustering is performed via a finite mixture of Mallows models. Inference on the number of clusters is then conducted a posteriori. Here we propose a Bayesian nonparametric Mallows model, based on a Dirichlet process mixture model. This allows joint inference on the number of non-empty clusters and on the clustering allocation, as well as posterior inference on cluster-specific parameters. The implementation of the proposed sampling algorithm is integrated into the existing R package BayesMallows, which also supports data in the form of incomplete rankings and pairwise comparisons. Simulated data show good performance of the nonparametric model compared to a finite mixture model in terms of recovery of the correct number of clusters, while empirical data on movie ratings show the model's effectiveness in providing personalized movie recommendations on discarded ratings.

stat.ME

Hyperevent network modelling of partially observed gossip data

Gossiping is a widespread social phenomenon that shapes relationships and information flow in communities. From a network theoretic point of view, gossiping can be seen as a higher-order interaction, as it involves at least two persons talking about a non-present third. The mechanism of gossiping is complex: it is most likely dynamic, as its intensity changes over time, and possibly viral, if a gossiping event induces future gossiping, such as a repetition or retaliation. We define covariates of interest for these effects and propose a relational hyperevent model to study and quantify these complex dynamics. We consider survey data collected yearly from 44 secondary schools in Hungary. No information is available about the exact timing of the events nor about the aggregate number of events within the yearly time interval. What is measured is whether at least one gossiping event has occurred in a given time interval. We extend inference for relational hyperevent models to the case of rightcensored interval-time data and show how flexible and efficient generalized additive models can be used for estimation of effects of interest. Our analysis on the school data illustrates how a model that accounts for linear, smooth and random effects can identify the social drivers of gossiping, while revealing complex temporal dynamics.

stat.ME

Causal generalized linear models via Pearson risk invariance

Prediction invariance of causal models under heterogeneous settings has been exploited by a number of recent methods for causal discovery, typically focussing on recovering the causal parents of a target variable of interest. Existing methods require observational data from a number of sufficiently different environments, which is rarely available. In this paper, we consider a structural equation model where the target variable is described by a generalized linear model conditional on its parents. Besides having finite moments, no modelling assumptions are made on the conditional distributions of the other variables in the system, and nonlinear effects on the target variable can naturally be accommodated by a generalized additive structure. Under this setting, we characterize the causal model uniquely by means of two key properties: the Pearson risk invariant under the causal model and, conditional on the causal parents, the causal parameters maximize the expected likelihood. These two properties form the basis of a computational strategy for searching the causal model among all possible models. A stepwise greedy search is proposed for systems with a large number of variables. Crucially, for generalized linear models with a known dispersion parameter, such as Poisson and logistic regression, the causal model can be identified from a single data environment. The method is implemented in the R package causalreg.

stat.ME

Causal drivers of dynamic networks

Dynamic networks models describe temporal interactions between social actors, and as such have been used to describe financial fraudulent transactions, dispersion of destructive invasive species across the globe, and the spread of fake news. An important question in all of these examples is what are the causal drivers underlying these processes. Current network models are exclusively descriptive and based on correlative structures. In this paper we propose a causal extension of dynamic network modelling. In particular, we prove that the causal model satisfies a set of population conditions that uniquely identifies the causal drivers. The empirical analogue of these conditions provide a consistent causal discovery algorithm, which distinguishes it from other inferential approaches. Crucially, data from a single environment is sufficient. We apply the method in an analysis of bike sharing data in Washington D.C. in July 2023.

stat.ME

Relational event models with global covariates

Bike sharing is an increasingly popular mobility choice as it is a sustainable, healthy and economically viable transportation mode. By interpreting rides between bike stations over time as temporal events connecting two bike stations, relational event models can provide important insights into this phenomenon. The focus of relational event models, as a typical event history model, is normally on dyadic or node-specific covariates, as global covariates are considered nuisance parameters in a partial likelihood approach. As full likelihood approaches are infeasible given the sheer size of the relational process, we propose an innovative sampling approach of temporally shifted non-events to recover important global drivers of the relational process. The method combines nested case-control sampling on a time-shifted version of the event process. This leads to a partial likelihood of the relational event process that is identical to that of a degenerate logistic additive model, enabling efficient estimation of both global and non-global covariate effects. The computational effectiveness of the method is demonstrated through a simulation study. The analysis of around 350,000 bike rides in the Washington D.C. area reveals significant influences of weather and time of day on bike sharing dynamics, besides a number of traditional node-specific and dyadic covariates.

stat.ME

Inferring the dynamics of quasi-reaction systems via nonlinear local mean-field approximations

In the modelling of stochastic phenomena, such as quasi-reaction systems, parameter estimation of kinetic rates can be challenging, particularly when the time gap between consecutive measurements is large. Local linear approximation approaches account for the stochasticity in the system but fail to capture the nonlinear nature of the underlying process. At the mean level, the dynamics of the system can be described by a system of ODEs, which have an explicit solution only for simple unitary systems. An analytical solution for generic quasi-reaction systems is proposed via a first order Taylor approximation of the hazard rate. This allows a nonlinear forward prediction of the future dynamics given the current state of the system. Predictions and corresponding observations are embedded in a nonlinear least-squares approach for parameter estimation. The performance of the algorithm is compared to existing SDE and ODE-based methods via a simulation study. Besides the increased computational efficiency of the approach, the results show an improvement in the kinetic rate estimation, particularly for data observed at large time intervals. Additionally, the availability of an explicit solution makes the method robust to stiffness, which is often present in biological systems. An illustration on Rhesus Macaque data shows the applicability of the approach to the study of cell differentiation.

stat.ME

A unified approach to penalized likelihood estimation of covariance matrices in high dimensions

We consider the problem of estimation of a covariance matrix for Gaussian data in a high dimensional setting. Existing approaches include maximum likelihood estimation under a pre-specified sparsity pattern, l_1-penalized loglikelihood optimization and ridge regularization of the sample covariance. We show that these three approaches can be addressed in an unified way, by considering the constrained optimization of an objective function that involves two suitably defined penalty terms. This unified procedure exploits the advantages of each individual approach, while bringing novelty in the combination of the three. We provide an efficient algorithm for the optimization of the regularized objective function and describe the relationship between the two penalty terms, thereby highlighting the importance of the joint application of the three methods. A simulation study shows how the sparse estimates of covariance matrices returned by the procedure are stable and accurate, both in low and high dimensional settings, and how their calculation is more efficient than existing approaches under a partially known sparsity pattern. An illustration on sonar data shows is presented for the identification of the covariance structure among signals bounced off a certain material. The method is implemented in the publicly available R package gicf.

stat.ME

Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems

High dimensional and heterogeneous count data are collected in various applied fields. In this paper, we look closely at high-resolution sequencing data on the microbiome, which have enabled researchers to study the genomes of entire microbial communities. Revealing the underlying interactions between these communities is of vital importance to learn how microbes influence human health. To perform structural learning from multivariate count data such as these, we develop a novel Gaussian copula graphical model with two key elements. Firstly, we employ parametric regression to characterize the marginal distributions. This step is crucial for accommodating the impact of external covariates. Neglecting this adjustment could potentially introduce distortions in the inference of the underlying network of dependences. Secondly, we advance a Bayesian structure learning framework, based on a computationally efficient search algorithm that is suited to high dimensionality. The approach returns simultaneous inference of the marginal effects and of the dependence structure, including graph uncertainty estimates. A simulation study and a real data analysis of microbiome data highlight the applicability of the proposed approach at inferring networks from multivariate count data in general, and its relevance to microbiome analyses in particular. The proposed method is implemented in the R package BDgraph.

stat.ME

Joint modelling of national cultures accounting for within and between-country heterogeneity

Cultural values vary significantly around the world. Despite a large heterogeneity, similarities across national cultures are present. This paper studies cross-country culture heterogeneity via the joint inference of country-specific copula graphical models from world-wide survey data. To this end, a random graph generative model of the cultural networks is introduced, with a latent space and proximity measures that embed cultural relatedness across countries. Within-country heterogeneity is also accounted for, via parametric modelling of the marginal distributions of each cultural trait. All together, the different components of the model are able to identify several dimensions of culture.

stat.ME

Latent event history models for quasi-reaction systems

Various processes can be modelled as quasi-reaction systems of stochastic differential equations, such as cell differentiation and disease spreading. Since the underlying data of particle interactions, such as reactions between proteins or contacts between people, are typically unobserved, statistical inference of the parameters driving these systems is developed from concentration data measuring each unit in the system over time. While observing the continuous time process at a time scale as fine as possible should in theory help with parameter estimation, the existing Local Linear Approximation (LLA) methods fail in this case, due to numerical instability caused by small changes of the system at successive time points. On the other hand, one may be able to reconstruct the underlying unobserved interactions from the observed count data. Motivated by this, we first formalise the latent event history model underlying the observed count process. We then propose a computationally efficient Expectation-Maximation algorithm for parameter estimation, with an extended Kalman filtering procedure for the prediction of the latent states. A simulation study shows the performance of the proposed method and highlights the settings where it is particularly advantageous compared to the existing LLA approaches. Finally, we present an illustration of the methodology on the spreading of the COVID-19 pandemic in Italy.

stat.ME

Random graphical model of microbiome interactions in related environments

The microbiome constitutes a complex microbial ecology of interacting components that regulates important pathways in the host. Measurements of microbial abundances are key to learning the intricate network of interactions amongst microbes. Microbial communities at various body sites tend to share some overall common structure, while also showing diversity related to the needs of the local environment. We propose a computational approach for the joint inference of microbiota systems from metagenomic data for a number of body sites. The random graphical model (RGM) allows for heterogeneity across the different environments while quantifying their relatedness at the structural level. In addition, the model allows for the inclusion of external covariates at both the microbial and interaction levels, further adapting to the richness and complexity of microbiome data. Our results show how: the RGM approach is able to capture varying levels of structural similarity across the different body sites and how this is supported by their taxonomical classification; the Bayesian implementation of the RGM fully quantifies parameter uncertainty; the microbiome network posteriors show not only a stable core, but also interesting individual differences between the various body sites, as well as interpretable relationships between various classes of microbes.

stat.ME

Cultures as networks of cultural traits: A unifying framework for measuring culture and cultural distances

Making use of the information from the World Value Survey (WVS), and operationalizing a definition of national culture that encompasses both the relevance of specific cultural traits and the interdependence among them, this paper proposes a methodology to reveal the latent structure of national culture and to measure cultural distance between countries that takes into account both the difference in cultural traits and the difference in the network structure of national cultures. Exploiting the possibilities offered by copula graphical models for discrete data, this paper infers the cultural networks of all the countries included in the WVS (Wave 6) and proposes a novel unifying framework to measure national culture and international cultural distances. The Jeffreys' divergence between copula graphical models, taken as the measure of cultural distance between countries, captures the orthogonality of the two components of cultural distance: the one based on cultural traits and the one based on the network structure among them. Moreover, the two components are shown to correlate with different national and structural characteristics of cultural networks, thus encompassing the different informational sets related to national cultures.

stat.AP

Sparse inference of the human hematopoietic system from heterogeneous and partially observed genomic data

Hematopoiesis is the process of blood cell formation, through which progenitor stem cells differentiate into mature forms, such as white and red blood cells or mature platelets. While the precursors of the mature forms share many regulatory pathways involving common cellular nuclear factors, specific networks of regulation shape their fate towards one lineage or another. In this study, we aim to analyse the complex regulatory network that drives the formation of mature red blood cells and platelets from their common precursor. To this aim, we develop a dedicated graphical model which we infer from the latest RT-qPCR genomic data. The model also accounts for the effect of external genomic data. A computationally efficient Expectation-Maximization algorithm allows regularised network inference from the high-dimensional and often only partially observed RT-qPCR data. A careful combination of alternating direction method of multipliers algorithms allows achieving sparsity in the individual lineage networks and a high sharing between these networks, together with the detection of the associations between the membrane-bound receptors and the nuclear factors. The approach will be implemented in the R package cglasso and can be used in similar applications where network inference is conducted from high-dimensional, heterogeneous and partially observed data.

stat.ME

Identifying overlapping terrorist cells from the Noordin Top actor-event network

Actor-event data are common in sociological settings, whereby one registers the pattern of attendance of a group of social actors to a number of events. We focus on 79 members of the Noordin Top terrorist network, who were monitored attending 45 events. The attendance or non-attendance of the terrorist to events defines the social fabric, such as group coherence and social communities. The aim of the analysis of such data is to learn about the affiliation structure. Actor-event data is often transformed to actor-actor data in order to be further analysed by network models, such as stochastic block models. This transformation and such analyses lead to a natural loss of information, particularly when one is interested in identifying, possibly overlapping, subgroups or communities of actors on the basis of their attendances to events. In this paper we propose an actor-event model for overlapping communities of terrorists, which simplifies interpretation of the network. We propose a mixture model with overlapping clusters for the analysis of the binary actor-event network data, called {\tt manet}, and develop a Bayesian procedure for inference. After a simulation study, we show how this analysis of the terrorist network has clear interpretative advantages over the more traditional approaches of affiliation network analysis.

stat.AP

Mixtures of multivariate generalized linear models with overlapping clusters

With the advent of ubiquitous monitoring and measurement protocols, studies have started to focus more and more on complex, multivariate and heterogeneous datasets. In such studies, multivariate response variables are drawn from a heterogeneous population often in the presence of additional covariate information. In order to deal with this intrinsic heterogeneity, regression analyses have to be clustered for different groups of units. Up until now, mixture model approaches assigned units to distinct and non-overlapping groups. However, not rarely these units exhibit more complex organization and clustering. It is our aim to define a mixture of generalized linear models with overlapping clusters of units. This involves crucially an overlap function, that maps the coefficients of the parent clusters into the the coefficient of the multiple allocation units. We present a computationally efficient MCMC scheme that samples the posterior distribution of the parameters in the model. An example on a two-mode network study shows details of the implementation in the case of a multivariate probit regression setting. A simulation study shows the overall performance of the method, whereas an illustration of the voting behaviour on the US supreme court shows how the 9 justices split in two overlapping sets of justices.

stat.ME