Searcharxiv⌕ Search

arXiv subjects

Peter X. K. Song

Publications and source records attributed to Peter X. K. Song.

12 recordsLinked to original sources

A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning

Data used to train predictive models via empirical risk minimization (ERM) often contain sensitive personal information. While differential privacy (DP) provides mathematically provable bounds to protect such data, previous work has focused almost exclusively on unweighted ERM. We consider weighted ERM (wERM) -- an important generalization where individual contributions to the objective function vary. We propose the first DP algorithm for general wERM with formal privacy guarantees and derive both its empirical and population utility bounds. Crucially, this general wERM framework provides a pathway for deriving privacy-preserving learning methods for individualized treatment rules, including the popular outcome-weighted learning (OWL) approach. We evaluate DP-wERM applied to OWL in simulated and real data experiments. Our empirical results demonstrate that training OWL models via wERM provides strong DP guarantees while maintaining robust performance, proving the method is practical for sensitive, real-world data.

stat.ML↗

Identification, Estimation, and Inference for Sequential Causally Ordered Mediation Pathways

Mediation analysis plays an essential role in uncovering the mechanisms by which an exposure influences an outcome through intermediate pathways. While methodological advances for single-mediator settings are well established, rigorous tools for handling multiple, sequentially ordered mediators remain underdeveloped. Such settings are common in applications like longitudinal cohort studies, where exposures operate through complex chains of mediators over time. In this paper, we establish a general framework for sequentially ordered mediators that enables the identification and formal decomposition of the total effect into component path-specific effects. We also develop estimation procedures for mediation estimands with both continuous and categorical outcomes. Furthermore, we introduce a new testing strategy to conduct inference using a studentized statistic combined with data-splitting. This approach achieves valid Type I error control under the composite null across diverse data-generating mechanisms. Through extensive simulations and applications to two large-scale empirical studies, we demonstrate that the proposed methodology provides reliable estimation, valid inference, and improved power for discovering novel mediation pathways.

stat.ME↗

Network Structural Equation Models for Causal Mediation and Spillover Effects

Social network interference induces complex dependencies where a unit's outcome is influenced not only by its own exposure and mediator but also by those of connected neighbors. In such settings, a significant challenge lies in distinguishing direct exposure effects from interference-driven spillover effects, and further separating these from indirect effects mediated by intermediate variables. To address this, we propose a theoretical framework utilizing structural graphical models. Central to our approach is the Random Effects Network Structural Equation Model (REN-SEM), which extends the exposure mapping paradigm to capture these multifaceted spillover and mediation mechanisms while accounting for latent dependencies within mediators and outcomes. We establish general identification conditions and derive decomposition formulas for six distinct mechanistic estimands. Furthermore, for the class of Linear REN-SEMs, we develop a maximum likelihood estimation framework and establish a rigorous asymptotic theory tailored to non-i.i.d. network data, proving the consistency of our estimators and the validity of the variance estimates. The robustness and practical utility of our methodology are demonstrated through simulation experiments and an analysis of the Twitch Gamers Network, underscoring its effectiveness in quantifying intricate network-mediated exposure effects.

stat.ME↗

Model-Assisted Causal Inference for the Treatment Effect on Recurrent Events in the Presence of Terminal Events

This paper is motivated by evaluating the benefits of patients receiving mechanical circulatory support (MCS) devices in end-stage heart failure management inference, in which hypothesis testing for a treatment effect on the risk of recurrent events is challenged in the presence of terminal events. Existing methods based on cumulative frequency unreasonably disadvantage longer survivors as they tend to experience more recurrent events. The While-Alive-based (WA) test has provided a solution to address this survival-length-bias problem, and it performs well when the recurrent event rate holds constant over time. However, if such a constant-rate assumption is violated, the WA test can exhibit an inflated type I error and inaccurate estimation of treatment effects. To fill this methodological gap, we propose a Proportional Rate Marginal Structural Model-assisted Test (PR-MSMaT) in the causal inference framework of separable treatment effects for recurrent and terminal events. Using the simulation study, we demonstrate that our PR-MSMaT can properly control type I error while gaining power comparable to the WA test under time-varying recurrent event rates. We employ PR-MSMaT to compare different MCS devices with the postoperative risk of gastrointestinal bleeding among patients enrolled in the Interagency Registry of Mechanically Assisted Circulatory Support program.

stat.AP↗

Post-2024 U.S. Presidential Election Analysis of Election and Poll Data: Real-life Validation of Prediction via Small Area Estimation and Uncertainty Quantification

We carry out a post-election analysis of the 2024 U.S. Presidential Election (USPE) using a prediction model derived from the Small Area Estimation (SAE) methodology. With pollster data obtained one week prior to the election day, retrospectively, our SAE-based prediction model can perfectly predict the Electoral College election results in all 44 states where polling data were available. In addition to such desirable prediction accuracy, we introduce the probability of incorrect prediction (PoIP) to rigorously analyze prediction uncertainty. Since the standard bootstrap method appears inadequate for estimating PoIP, we propose a conformal inference method that yields reliable uncertainty quantification. We further investigate potential pollster biases by the means of sensitivity analyses and conclude that swing states are particularly vulnerable to polling bias in the prediction of the 2024 USPE.

stat.AP↗

Asymmetric predictability in causal discovery: an information theoretic approach

Causal investigations in observational studies pose a great challenge in research where randomized trials or intervention-based studies are not feasible. We develop an information geometric causal discovery and inference framework of "predictive asymmetry". For $(X, Y)$, predictive asymmetry enables assessment of whether $X$ is more likely to cause $Y$ or vice-versa. The asymmetry between cause and effect becomes particularly simple if $X$ and $Y$ are deterministically related. We propose a new metric called the Directed Mutual Information ($DMI$) and establish its key statistical properties. $DMI$ is not only able to detect complex non-linear association patterns in bivariate data, but also is able to detect and infer causal relations. Our proposed methodology relies on scalable non-parametric density estimation using Fourier transform. The resulting estimation method is manyfold faster than the classical bandwidth-based density estimation. We investigate key asymptotic properties of the $DMI$ methodology and a data-splitting technique is utilized to facilitate causal inference using the $DMI$. Through simulation studies and an application, we illustrate the performance of $DMI$.

stat.ME↗

fastMI: a fast and consistent copula-based estimator of mutual information

As a fundamental concept in information theory, mutual information ($MI$) has been commonly applied to quantify association between random vectors. Most existing nonparametric estimators of $MI$ have unstable statistical performance since they involve parameter tuning. We develop a consistent and powerful estimator, called fastMI, that does not incur any parameter tuning. Based on a copula formulation, fastMI estimates $MI$ by leveraging Fast Fourier transform-based estimation of the underlying density. Extensive simulation studies reveal that fastMI outperforms state-of-the-art estimators with improved estimation accuracy and reduced run time for large data sets. fastMI provides a powerful test for independence that exhibits satisfactory type I error control. Anticipating that it will be a powerful tool in estimating mutual information in a broad range of data, we develop an R package fastMI for broader dissemination.

stat.AP↗

Robust High-Dimensional Regression with Coefficient Thresholding and its Application to Imaging Data Analysis

It is of importance to develop statistical techniques to analyze high-dimensional data in the presence of both complex dependence and possible outliers in real-world applications such as imaging data analyses. We propose a new robust high-dimensional regression with coefficient thresholding, in which an efficient nonconvex estimation procedure is proposed through a thresholding function and the robust Huber loss. The proposed regularization method accounts for complex dependence structures in predictors and is robust against outliers in outcomes. Theoretically, we analyze rigorously the landscape of the population and empirical risk functions for the proposed method. The fine landscape enables us to establish both {statistical consistency and computational convergence} under the high-dimensional setting. The finite-sample properties of the proposed method are examined by extensive simulation studies. An illustration of real-world application concerns a scalar-on-image regression analysis for an association of psychiatric disorder measured by the general factor of psychopathology with features extracted from the task functional magnetic resonance imaging data in the Adolescent Brain Cognitive Development study.

stat.ME↗

Data Discovery Using Lossless Compression-Based Sparse Representation

Sparse representation has been widely used in data compression, signal and image denoising, dimensionality reduction and computer vision. While overcomplete dictionaries are required for sparse representation of multidimensional data, orthogonal bases represent one-dimensional data well. In this paper, we propose a data-driven sparse representation using orthonormal bases under the lossless compression constraint. We show that imposing such constraint under the Minimum Description Length (MDL) principle leads to a unique and optimal sparse representation for one-dimensional data, which results in discriminative features useful for data discovery.

eess.SP↗

Adaptive multi-channel event segmentation and feature extraction for monitoring health outcomes

$\textbf{Objective}$: To develop a multi-channel device event segmentation and feature extraction algorithm that is robust to changes in data distribution. $\textbf{Methods}$: We introduce an adaptive transfer learning algorithm to classify and segment events from non-stationary multi-channel temporal data. Using a multivariate hidden Markov model (HMM) and Fisher's linear discriminant analysis (FLDA) the algorithm adaptively adjusts to shifts in distribution over time. The proposed algorithm is unsupervised and learns to label events without requiring $\textit{a priori}$ information about true event states. The procedure is illustrated on experimental data collected from a cohort in a human viral challenge (HVC) study, where certain subjects have disrupted wake and sleep patterns after exposure to a H1N1 influenza pathogen. $\textbf{Results}$: Simulations establish that the proposed adaptive algorithm significantly outperforms other event classification methods. When applied to early time points in the HVC data the algorithm extracts sleep/wake features that are predictive of both infection and infection onset time. $\textbf{Conclusion}$: The proposed transfer learning event segmentation method is robust to temporal shifts in data distribution and can be used to produce highly discriminative event-labeled features for health monitoring. $\textbf{Significance}$: Our integrated multisensor signal processing and transfer learning method is applicable to many ambulatory monitoring applications.

eess.SP↗

Pattern-Based Analysis of Time Series: Estimation

While Internet of Things (IoT) devices and sensors create continuous streams of information, Big Data infrastructures are deemed to handle the influx of data in real-time. One type of such a continuous stream of information is time series data. Due to the richness of information in time series and inadequacy of summary statistics to encapsulate structures and patterns in such data, development of new approaches to learn time series is of interest. In this paper, we propose a novel method, called pattern tree, to learn patterns in the times-series using a binary-structured tree. While a pattern tree can be used for many purposes such as lossless compression, prediction and anomaly detection, in this paper we focus on its application in time series estimation and forecasting. In comparison to other methods, our proposed pattern tree method improves the mean squared error of estimation.

stat.ME↗

An unsupervised transfer learning algorithm for sleep monitoring

Objective: To develop multisensor-wearable-device sleep monitoring algorithms that are robust to health disruptions affecting sleep patterns. Methods: We develop an unsupervised transfer learning algorithm based on a multivariate hidden Markov model and Fisher's linear discriminant analysis, adaptively adjusting to sleep pattern shift by training on dynamics of sleep/wake states. The proposed algorithm operates, without requiring a priori information about true sleep/wake states, by establishing an initial training set with hidden Markov model and leveraging a taper window mechanism to learn the sleep pattern in an incremental fashion. Our domain-adaptation algorithm is applied to a dataset collected in a human viral challenge study to identify sleep/wake periods of both uninfected and infected participants. Results: The algorithm successfully detects sleep/wake sessions in subjects whose sleep patterns are disrupted by respiratory infection (H3N2 flu virus). Pre-symptomatic features based on the detected periods are found to be strongly predictive of both infection status (AUC = 0.844) and infection onset time (AUC = 0.885), indicating the effectiveness and usefulness of the algorithm. Conclusion: Our method can effectively detect sleep/wake states in the presence of sleep pattern shift. Significance: Utilizing integrated multisensor signal processing and adaptive training schemes, our algorithm is able to capture key sleep patterns in ambulatory monitoring, leading to better automated sleep assessment and prediction.

stat.AP↗