Searcharxiv⌕ Search

arXiv subjects

Victor De Gruttola

Publications and source records attributed to Victor De Gruttola.

10 recordsLinked to original sources

CCMnet: A Software Package for Network Generation with Congruence Class Models

We introduce CCMnet, an R package designed to generate network ensembles that accurately reflect the uncertainty inherent in empirical data. While traditional network modeling often results in ensembles with fixed property values or model-determined levels of variability, CCMnet enables a continuous spectrum of variability for network properties, including edge counts, degree distribution, and mixing patterns. By defining probability distributions directly over congruence classes of networks, the package allows researchers to specify the uncertainty in network properties across the generated ensemble to match a specific sampling design or empirical distribution. Furthermore, this formulation provides a principled framework that encompasses several classic models (e.g., Erdős--Rényi model, stochastic block models, and certain exponential random graph models) that implicitly share this structural basis, while offering the flexibility to specify arbitrary, even non-parametric, distributions for network properties. CCMnet implements a Markov chain Monte Carlo (MCMC) framework to sample from these models. The utility of the package is illustrated by generating posterior predictive network ensembles representing school friendship networks.

stat.CO↗

Investigating the HIV Epidemic in Miami Using a Novel Approach for Bayesian Inference on Partially Observed Networks

Molecular HIV Surveillance (MHS) has been described as key to enabling rapid responses to HIV outbreaks. It operates by linking individuals with genetically similar viral sequences, which forms a network. A major limitation of MHS is that it depends on sequence collection, which very rarely covers the entire population of interest. Ignoring missing data by conducting complete case analysis--which assumes that the observed network is complete--has been shown to result in significantly biased estimates of network properties. We use MHS to investigate disease dynamics of the HIV epidemic in Miami-Dade County (MDC) among men who have sex with men (MSM)--only 30.1% have a reported sequence. To do so, we present an approach for making Bayesian inferences on partially observed networks. Through a simulation study, we demonstrate a reduction in error of 43%-63% between our estimates and complete case analyses. We estimate increased mixing between MSM communities in MDC, defined by race and transmission risk compared to the results based on complete case analysis. Our approach makes use of a flexible network model--congruence class model--to overcome the high computational burden of previously reported Bayesian approaches to estimate network properties from partially observed networks.

stat.AP↗

Leveraging Contact Network Information in Clustered Randomized Studies of Contagion Processes

In a randomized study, leveraging covariates related to the outcome (e.g. disease status) may produce less variable estimates of the effect of exposure. For contagion processes operating on a contact network, transmission can only occur through ties that connect affected and unaffected individuals; the outcome of such a process is known to depend intimately on the structure of the network. In this paper, we investigate the use of contact network features as efficiency covariates in exposure effect estimation. Using augmented generalized estimating equations (GEE), we estimate how gains in efficiency depend on the network structure and spread of the contagious agent or behavior. We apply this approach to simulated randomized trials using a stochastic compartmental contagion model on a collection of model-based contact networks and compare the bias, power, and variance of the estimated exposure effects using an assortment of network covariate adjustment strategies. We also demonstrate the use of network-augmented GEEs on a clustered randomized trial evaluating the effects of wastewater monitoring on COVID-19 cases in residential buildings at the the University of California San Diego.

stat.AP↗

Estimating Viral Genetic Linkage Rates in the Presence of Missing Data

Although the interest in the the use of social and information networks has grown, most inferences on networks assume the data collected represents the complete. However, when ignoring missing data, even when missing completely at random, this results in bias for estimators regarding inference network related parameters. In this paper, we focus on constructing estimators for the probability that a randomly selected node has node has at least one edge under the assumption that nodes are missing completely at random along with their corresponding edges. In addition, issues also arise in obtaining asymptotic properties for such estimators, because linkage indicators across nodes are correlated preventing the direct application of the Central Limit Theorem and Law of Large Numbers. Using a subsampling approach, we present an improved estimator for our parameter of interest that accommodates for missing data. Utilizing the theory U-statistics, we derive consistency and asymptotic normality of the proposed estimator. This approach decreases the bias in estimating our parameter of interest. We illustrate our approach using the HIV viral strains from a large cluster-randomized trial of a combination HIV prevention intervention -- the Botswana Combination Prevention Project (BCPP).

stat.ME↗

Investigation of Patient-sharing Networks Using a Bayesian Network Model Selection Approach for Congruence Class Models

A Bayesian approach to conduct network model selection is presented for a general class of network models referred to as the congruence class models (CCMs). CCMs form a broad class that includes as special cases several common network models, such as the Erdős-Rényi-Gilbert model, stochastic block model and many exponential random graph models. Due to the range of models able to be specified as a CCM, investigators are better able to select a model consistent with generative mechanisms associated with the observed network compared to current approaches. In addition, the approach allows for incorporation of prior information. We utilize the proposed Bayesian network model selection approach for CCMs to investigate several mechanisms that may be responsible for the structure of patient-sharing networks, which are associated with the cost and quality of medical care. We found evidence in support of heterogeneity in sociality but not selective mixing by provider type nor degree.

stat.AP↗

Recursive Formula for Labeled Graph Enumeration

This manuscript presents a general recursive formula to estimate the size of fibers associated with algebraic maps from graphs to summary statistics of importance for social network analysis, such as number of edges (graph density), degree sequence, degree distribution, mixing by nodal covariates, and degree mixing. That is, the formula estimates the number of labeled graphs that have given values for network properties. The proposed approach can be extended to additional network properties (e.g., clustering) as well as properties of bipartite networks. For special settings in which alternative formulas exist, simulation studies demonstrate the validity of the proposed approach. We illustrate the approach for estimating the size of fibers associated with the Barabási--Albert model for the properties of degree distribution and degree mixing. In addition, we demonstrate how the approach can be used to assess the diversity of graphs within a fiber.

stat.CO↗

Dynamic Network Prediction

We present a statistical framework for generating predicted dynamic networks based on the observed evolution of social relationships in a population. The framework includes a novel and flexible procedure to sample dynamic networks given a probability distribution on evolving network properties; it permits the use of a broad class of approaches to model trends, seasonal variability, uncertainty, and changes in population composition. Current methods do not account for the variability in the observed historical networks when predicting the network structure; the proposed method provides a principled approach to incorporate uncertainty in prediction. This advance aids in the designing of network-based interventions, as development of such interventions often requires prediction of the network structure in the presence and absence of the intervention. Two simulation studies are conducted to demonstrate the usefulness of generating predicted networks when designing network-based interventions. The framework is also illustrated by investigating results of potential interventions on bill passage rates using a dynamic network that represents the sponsor/co-sponsor relationships among senators derived from bills introduced in the US Senate from 2003-2016.

cs.SI↗

Novel Methods for the Analysis of Stepped Wedge Cluster Randomized Trials

Stepped wedge cluster randomized trials (SW-CRTs) have become increasingly popular and are used for a variety of interventions and outcomes, often chosen for their feasibility advantages. SW-CRTs must account for time trends in the outcome because of the staggered rollout of the intervention inherent in the design. Robust inference procedures and non-parametric analysis methods have recently been proposed to handle such trends without requiring strong parametric modeling assumptions, but these are less powerful than model-based approaches. We propose several novel analysis methods that reduce reliance on modeling assumptions while preserving some of the increased power provided by the use of mixed effects models. In one method, we use the synthetic control approach to find the best matching clusters for a given intervention cluster. This approach can improve the power of the analysis but is fully non-parametric. Another method makes use of within-cluster crossover information to construct an overall estimator. We also consider methods that combine these approaches to further improve power. We test these methods on simulated SW-CRTs and identify settings for which these methods gain robustness to model misspecification while retaining some of the power advantages of mixed effects models. Finally, we propose avenues for future research on the use of these methods; motivation for such research arises from their flexibility, which allows the identification of specific causal contrasts of interest, their robustness, and the potential for incorporating covariates to further increase power. Investigators conducting SW-CRTs might well consider such methods when common modeling assumptions may not hold.

stat.ME↗

Leveraging contact network structure in the design of cluster randomized trials

Background: In settings where proof-of-principle trials have succeeded but the effectiveness of different forms of implementation remains uncertain, trials that not only generate information about intervention effects but also provide public health benefit would be useful. Cluster randomized trials (CRT) capture both direct and indirect intervention effects; the latter depends heavily on contact networks within and across clusters. We propose a novel class of connectivity-informed trial designs that leverages information about such networks in order to improve public health impact and preserve ability to detect intervention effects. Methods: We consider CRTs in which the order of enrollment is based on the total number of ties between individuals across clusters (based either on the total number of inter-cluster connections or on connections only to untreated clusters). We include options analogous both to traditional Parallel and Stepped Wedge designs. We also allow for control clusters to be "held-back" from re-randomization for some period. We investigate the performance epidemic control and power to detect vaccine effect performance of these designs by simulating vaccination trials during an SEIR-type epidemic using a network-structured agent-based model. Results: In our simulations, connectivity-informed designs have lower peak infectiousness than comparable traditional designs and reduce cumulative incidence by 20%, but with little impact on time to end of epidemic and reduced power to detect differences in incidence across clusters. However even a brief "holdback" period restores most of the power lost compared to traditional approaches. Conclusion: Incorporating information about cluster connectivity in design of CRTs can increase their public health impact, especially in acute outbreak settings, with modest cost in power to detect an effective intervention.

q-bio.QM↗

Flexible covariate-adjusted exact tests of randomized treatment effects with application to a trial of HIV education

The primary goal of randomized trials is to compare the effects of different interventions on some outcome of interest. In addition to the treatment assignment and outcome, data on baseline covariates, such as demographic characteristics or biomarker measurements, are typically collected. Incorporating such auxiliary covariates in the analysis of randomized trials can increase power, but questions remain about how to preserve type I error when incorporating such covariates in a flexible way, particularly when the number of randomized units is small. Using the Young Citizens study, a cluster-randomized trial of an educational intervention to promote HIV awareness, we compare several methods to evaluate intervention effects when baseline covariates are incorporated adaptively. To ascertain the validity of the methods shown in small samples, extensive simulation studies were conducted. We demonstrate that randomization inference preserves type I error under model selection while tests based on asymptotic theory may yield invalid results. We also demonstrate that covariate adjustment generally increases power, except at extremely small sample sizes using liberal selection procedures. Although shown within the context of HIV prevention research, our conclusions have important implications for maximizing efficiency and robustness in randomized trials with small samples across disciplines.

stat.AP↗