SearcharxivSearch

arXiv subjects

Yuhua Zhang

Publications and source records attributed to Yuhua Zhang.

7 recordsLinked to original sources

A case study of causal mediation using Bayesian nonparametrics and semiparametric corrections

We propose a Bayesian nonparametric approach using a truncated Enriched Dirichlet Process mixture (EDPM) model to estimate natural direct (NDE) and indirect (NIE) effects in causal mediation analyses in the presence of post-treatment confounders. We introduce an efficient cluster reallocation Metropolis-Hasting algorithm to improve mixing in the blocked Gibbs sampler. We implement a one-step posterior correction based on the efficient influence function for our setting. This post-processing step solves a critical problem in Bayesian nonparametrics: how to obtain reliable estimates and posteriors for a specific causal estimand of interest (the NDE and NIE) with excellent frequentist properties, such as correct coverage, from a model designed for complex joint distributions. We conduct simulation studies to assess our method's performance and apply it to evaluate causal mediation effects in a weight management clinical trial.

stat.ME

Covariate Selection for Joint Latent Space Modeling of Sparse Network Data

Network data are increasingly common in the social sciences and infectious disease epidemiology. Analyses often link network structure to node-level covariates, but existing methods falter with sparse networks and high-dimensional node features. We propose a joint latent space modeling framework for sparse networks with high-dimensional binary node covariates that performs covariate selection while accounting for uncertainty in estimated latent positions. Building on joint latent space models that couple edges and node variables through shared latent positions, we introduce a group lasso screening step and incorporate a measurement-error-aware stabilization term to mitigate bias from using estimated latent positions as predictors. We establish prediction error rates for the covariate component both when latent positions are treated as observed and when they are estimated with bounded error; under uniform control across $q$ covariates and $n$ nodes, the rate is of order $O(\log q / n)$ up to an additional term due to latent position estimation error. Our method addresses three challenges: (1) incorporating information from isolated nodes, which are common in sparse networks but often ignored; (2) selecting relevant covariates from high-dimensional spaces; and (3) accounting for uncertainty in estimated latent positions. Simulations show predictive performance remains stable as covariate sparsity grows, while naive approaches degrade. We illustrate how the method can support efficient study design using household social networks from 75 Indian villages, where an emulated pilot study screens a large covariate battery and substantially reduces required subsequent data collection without sacrificing network predictive accuracy.

stat.ME

Identification and Estimation of Heterogeneous Interference Effects under Unknown Network

Interference--in which a unit's outcome is affected by the treatment of other units--poses significant challenges for the identification and estimation of causal effects. Most existing methods for estimating interference effects assume that the interference networks are known. In many practical settings, this assumption is unrealistic as such networks are typically latent. To address this challenge, we propose a novel framework for identifying and estimating heterogeneous group-level interference effects without requiring a known interference network. Specifically, we assume a shared latent community structure between the observed network and the unknown interference network. We demonstrate that interference effects are identifiable if and only if group-level interference effects are heterogeneous, and we establish the consistency and asymptotic normality of the maximum likelihood estimator (MLE). To handle the intractable likelihood function and facilitate the computation, we propose a Bayesian implementation and show that the posterior concentrates around the MLE. A series of simulation studies demonstrate the effectiveness of the proposed method and its superior performance compared with competitors. We apply our proposed framework to the encounter data of stroke patients from the California Department of Healthcare Access and Information (HCAI) and evaluate the causal interference effects of certain intervention in one hospital on the outcomes of other hospitals.

stat.ME

Community Detection through Recursive Partitioning in Bayesian Framework

Community detection involves grouping the nodes in the network and is one of the most-studied tasks in network science. Conventional methods usually require the specification of the number of communities $K$ in the network. This number is determined heuristically or by certain model selection criteria. In practice, different model selection criteria yield different values of $K$, leading to different results. We propose a community detection method based on recursive partitioning within the Bayesian framework. The method is compatible with a wide range of existing model-based community detection frameworks. In particular, our method does not require pre-specification of the number of communities and can capture the hierarchical structure of the network. We establish the theoretical guarantee of consistency under the stochastic block model and demonstrate the effectiveness of our method through simulations using different models that cover a broad range of scenarios. We apply our method to the California Department of Healthcare Access and Information (HCAI) data, including all Emergency Department (ED) and hospital discharges from 342 hospitals to identify regional hospital clusters.

stat.ME

AMANet: Advancing SAR Ship Detection with Adaptive Multi-Hierarchical Attention Network

Recently, methods based on deep learning have been successfully applied to ship detection for synthetic aperture radar (SAR) images. Despite the development of numerous ship detection methodologies, detecting small and coastal ships remains a significant challenge due to the limited features and clutter in coastal environments. For that, a novel adaptive multi-hierarchical attention module (AMAM) is proposed to learn multi-scale features and adaptively aggregate salient features from various feature layers, even in complex environments. Specifically, we first fuse information from adjacent feature layers to enhance the detection of smaller targets, thereby achieving multi-scale feature enhancement. Then, to filter out the adverse effects of complex backgrounds, we dissect the previously fused multi-level features on the channel, individually excavate the salient regions, and adaptively amalgamate features originating from different channels. Thirdly, we present a novel adaptive multi-hierarchical attention network (AMANet) by embedding the AMAM between the backbone network and the feature pyramid network (FPN). Besides, the AMAM can be readily inserted between different frameworks to improve object detection. Lastly, extensive experiments on two large-scale SAR ship detection datasets demonstrate that our AMANet method is superior to state-of-the-art methods.

cs.CV

Node-level community detection within edge exchangeable models for interaction processes

Scientists are increasingly interested in discovering community structure from modern relational data arising on large-scale social networks. While many methods have been proposed for learning community structure, few account for the fact that these modern networks arise from processes of interactions in the population. We introduce block edge exchangeable models (BEEM) for the study of interaction networks with latent node-level community structure. The block vertex components model (B-VCM) is derived as a canonical example. Several theoretical and practical advantages over traditional vertex-centric approaches are highlighted. In particular, BEEMs allow for sparse degree structure and power-law degree distributions within communities. Our theoretical analysis bounds the misspecification rate of block assignments, while supporting simulations show the properties of the network can be recovered. A computationally tractable Gibbs algorithm is derived. We demonstrate the proposed model using post-comment interaction data from Talklife, a large-scale online peer-to-peer support network, and contrast the learned communities from those using standard algorithms including spectral clustering and degree-correct stochastic block models.

stat.ME

D2D Big Data: Content Deliveries over Wireless Device-to-Device Sharing in Large Scale Mobile Networks

Recently the topic of how to effectively offload cellular traffic onto device-to-device (D2D) sharing among users in proximity has been gaining more and more attention of global researchers and engineers. Users utilize wireless short-range D2D communications for sharing contents locally, due to not only the rapid sharing experience and free cost, but also high accuracy on deliveries of interesting and popular contents, as well as strong social impacts among friends. Nevertheless, the existing related studies are mostly confined to small-scale datasets, limited dimensions of user features, or unrealistic assumptions and hypotheses on user behaviors. In this article, driven by emerging Big Data techniques, we propose to design a big data platform, named D2D Big Data, in order to encourage the wireless D2D communications among users effectively, to promote contents for providers accurately, and to carry out offloading intelligence for operators efficiently. We deploy a big data platform and further utilize a large-scale dataset (3.56 TBytes) from a popular D2D sharing application (APP), which contains 866 million D2D sharing activities on 4.5 million files disseminated via nearly 850 million users in 13 weeks. By abstracting and analyzing multidimensional features, including online behaviors, content properties, location relations, structural characteristics, meeting dynamics, social arborescence, privacy preservation policies and so on, we verify and evaluate the D2D Big Data platform regarding predictive content propagating coverage. Finally, we discuss challenges and opportunities regarding D2D Big Data and propose to unveil a promising upcoming future of wireless D2D communications.

cs.NI