SearcharxivSearch

arXiv subjects

Yuanjia Wang

Publications and source records attributed to Yuanjia Wang.

At least 19 recordsLinked to original sources

Debiased Inference for High-Dimensional Regression Models Based on Profile M-Estimation

Debiased inference for high-dimensional regression models has received substantial recent attention to ensure regularized estimators have valid inference. Many existing methods focus on achieving Neyman orthogonality through explicitly constructing projections onto the space of nuisance parameters, which is infeasible when an explicit form of the projection is unavailable. We introduce a general debiasing framework, Debiased Profile $M$-Estimation (DPME), which applies to a broad class of models and does not require model-specific Neyman orthogonalization or projection derivations as in existing methods. Our approach begins with obtaining an initial estimator of the parameters by optimizing a penalized objective function. To correct for the bias introduced by penalization, we construct a one-step estimator using the Newton--Raphson update, applied to the gradient of a profile function defined as the optimal objective function with the parameter of interest held fixed. We use numerical differentiation without requiring explicit calculation of the gradients. The resulting DPME estimator is shown to be asymptotically linear and normally distributed. Through extensive simulations, we demonstrate that the proposed method achieves better coverage rates than existing alternatives with largely reduced computational cost. Finally, we illustrate the utility of our method by applying it to estimate a treatment rule for multiple myeloma.

stat.ME

Integrative learning of individualized treatment rules from multiple studies with partially overlapping treatments

An individualized treatment rule (ITR) tailors treatments to a patient's specific characteristics. However, randomized controlled trials (RCTs) are often underpowered to detect the treatment effect heterogeneity needed for reliable ITR estimation. To address this limitation, there is growing interest in leveraging information from multiple studies to improve statistical power and support individualized decision-making. A key challenge in this context is that available RCTs may not evaluate the same set of treatments. In this paper, we propose an integrative learning framework that synthesizes evidence across multiple RCTs that share a common comparator but differ in their alternative treatment arms. Our method integrates information through a regularized weighted misclassification risk function and adaptively determines the contribution of each study to the ITRs of the others. We rigorously study the excess risk of the resulting estimator. Simulation studies demonstrate that the proposed approaches improve the estimation of both value and benefit functions. We illustrate the utility of our methodology using data from two landmark studies of major depressive disorder: the Establishing Moderators and Biosignatures of Antidepressant Response in Clinical Care study and the International Study to Predict Optimized Treatment in Depression study, both of which include a selective serotonin reuptake inhibitor as a common treatment arm. We find that the separate learning method outperforms one-size-fits-all methods, and our integrative methods further improve performance.

stat.ME

Shared hidden-factor information framework for multiple behavioral tasks

Understanding cognitive processes in major depressive disorder (MDD) often relies on behavioral tasks, which are typically analyzed separately, overlooking potential correlations and shared latent structure. To address this limitation, we propose the Shared Hidden-factor Information Framework for Multiple Behavioral Tasks (SHIFT), a joint modeling approach that leverages shared information across tasks, allowing each task to benefit from information learned by the others. SHIFT introduces subject-specific latent factors that capture cross-task dependencies while accommodating individual heterogeneity in decision-making, response times (RTs), and strategy switching. To address computational challenges without requiring high-dimensional integration, we develop an expectation-maximization with variational approximation algorithm that preserves both temporal structure and between-task dependencies. Through extensive simulation studies, we demonstrate that SHIFT substantially improves estimation accuracy and efficiency relative to single-task analyses. We then apply SHIFT to a study of MDD to jointly model the Probabilistic Reward Task (PRT) and the Flanker Task (FT). Results indicate that MDD participants show lower engagement in the PRT and reduced focus in the FT compared with healthy controls. Moreover, when individuals are engaged and focused, they exhibit longer RTs. Although observed RTs do not predict treatment response, the shared parameters recovered by SHIFT showed suggestive treatment-modulation patterns, indicating their potential as exploratory behavioral markers for therapeutic outcomes.

stat.ME

Maximin Learning of Individualized Treatment Effect on Multi-Domain Outcomes

Precision mental health requires treatment decisions that account for heterogeneous symptoms reflecting multiple clinical domains. However, existing methods for estimating individualized treatment effects (ITE) rely on a single summary outcome or a specific set of observed symptoms or measures, which are sensitive to symptom selection and limit generalizability to unmeasured yet clinically relevant domains. We propose DRIFT, a new maximin framework for estimating robust ITEs from high-dimensional item-level data by leveraging latent factor representations and adversarial learning. DRIFT learns latent constructs via generalized factor analysis, then constructs an anchored on-target uncertainty set that extrapolates beyond the observed measures to approximate the broader hyper-population of potential outcomes. By optimizing worst-case performance over this uncertainty set, DRIFT yields ITEs that are robust to underrepresented or unmeasured domains. We further show that DRIFT is invariant to admissible reparameterizations of the latent factors and admits a closed-form maximin solution, with theoretical guarantees for identification and convergence. In analyses of a randomized controlled trial for major depressive disorder (EMBARC), DRIFT demonstrates superior performance and improved generalizability to external multi-domain outcomes, including side effects and self-reported symptoms not used during training.

cs.LG

Cumulative Marginal Mean Model for Assessing Sequential Effects Using Digital Health Data

Mobile health (mHealth) leverages digital technologies, such as mobile phones, to capture objective, frequent, and real-world digital phenotypes from individuals, enabling the delivery of tailored interventions to accommodate substantial between-subject and temporal heterogeneity. However, evaluating heterogeneous treatment effects (HTEs) using digital phenotype data is challenging because treatments are delivered dynamically over time and may generate carryover effects that persist beyond the immediate response. Additionally, modeling observational data is complicated by confounding factors. To address these challenges, we propose a double machine learning (DML) method for estimating time-varying HTEs using digital phenotypes under a cumulative marginal mean model that separates current instantaneous effects from lagged carryover effects. Our approach uses a sequential estimation procedure together with Neyman-orthogonal scores to obtain robust inference for the time-varying HTEs. We establish the asymptotic normality of the proposed estimator. Extensive simulation studies validate the finite-sample performance of our approach, demonstrating the advantages of DML and the decomposition of treatment effects. We apply the method to an mHealth study of Parkinson's disease (PD), where we find that treatment is significantly more effective for younger patients. Our results highlight the potential of the proposed approach for advancing precision medicine in mHealth studies.

stat.ME

Adaptive Loss-tolerant Syndrome Measurements

In the presence of qubit losses, the building blocks of fault-tolerant error correction (FTEC) must be revisited. Existing loss-tolerant approaches are mainly architecture-specific, and little attention has been given to optimizing the syndrome measurement sequences under loss. Schemes designed for the standard Pauli error model are not directly applicable because the syndrome patterns differ when both Pauli errors and erasures can occur. Based on recent advances in loss detection units and loss-tolerant syndrome extraction gadgets, we extend the study of adaptive Shor-style measurement sequences to the mixed error model. We begin by discussing how to adaptively convert correctable erasures into located errors. The minimal overhead is quantified by the number of stabilizer measurements, which can be reduced to a subgroup dimension problem for erasures arising in any FTEC circuit for qubits and prime-dimensional qudits. As a byproduct, we provide the construction of the canonical generating set with respect to a given bipartite partition for a stabilizer group on qudits of composite dimension. We then generalize both the weak and strong FTEC conditions. Finally, we present adaptive syndrome-measurement protocols for the mixed error model, generalizing the adaptive protocols for the standard Pauli error model.

quant-ph

Semiparametric Analysis of Interval-Censored Data Subject to Inaccurate Diagnoses with A Terminal Event

Interval-censoring frequently occurs in studies of chronic diseases where disease status is inferred from intermittently collected biomarkers. Although many methods have been developed to analyze such data, they typically assume perfect disease diagnosis, which often does not hold in practice due to the inherent imperfect clinical diagnosis of cognitive functions or measurement errors of biomarkers such as cerebrospinal fluid. In this work, we introduce a semiparametric modeling framework using the Cox proportional hazards model to address interval-censored data in the presence of inaccurate disease diagnosis. Our model incorporates sensitivity and specificity of the diagnosis to account for uncertainty in whether the interval truly contains the disease onset. Furthermore, the framework accommodates scenarios involving a terminal event and when diagnosis is accurate, such as through postmortem analysis. We propose a nonparametric maximum likelihood estimation method for inference and develop an efficient EM algorithm to ensure computational feasibility. The regression coefficient estimators are shown to be asymptotically normal, achieving semiparametric efficiency bounds. We further validate our approach through extensive simulation studies and an application assessing Alzheimer's disease (AD) risk. We find that amyloid-beta is significantly associated with AD, but Tau is predictive of both AD and mortality.

stat.ME

Multi-level Latent Variable Models for Coheritability Analysis in Electronic Health Records

Electronic health records (EHRs) linked with familial relationship data offer a unique opportunity to investigate the genetic architecture of complex phenotypes at scale. However, existing heritability and coheritability estimation methods often fail to account for the intricacies of familial correlation structures, heterogeneity across phenotype types, and computational scalability. We propose a robust and flexible statistical framework for jointly estimating heritability and genetic correlation among continuous and binary phenotypes in EHR-based family studies. Our approach builds on multi-level latent variable models to decompose phenotypic covariance into interpretable genetic and environmental components, incorporating both within- and between-family variations. We derive iteration algorithms based on generalized equation estimations (GEE) for estimation. Simulation studies under various parameter configurations demonstrate that our estimators are consistent and yield valid inference across a range of realistic settings. Applying our methods to real-world EHR data from a large, urban health system, we identify significant genetic correlations between mental health conditions and endocrine/metabolic phenotypes, supporting hypotheses of shared etiology. This work provides a scalable and rigorous framework for coheritability analysis in high-dimensional EHR data and facilitates the identification of shared genetic influences in complex disease networks.

stat.ME

An Integrative Approach for Subtyping Mental Disorders Using Multimodal Data

Understanding the biological and behavioral heterogeneity underlying psychiatric disorders is critical for advancing precision diagnosis, treatment, and prevention. This paper addresses the scientific question of how multimodal data, spanning clinical, cognitive, and neuroimaging measures, can be integrated to identify biologically meaningful subtypes of mental disorders. We introduce Mixed INtegrative Data Subtyping (MINDS), a Bayesian hierarchical model designed to jointly analyze mixed-type data for simultaneous dimension reduction and clustering. Using data from the Adolescent Brain Cognitive Development (ABCD) Study, MINDS integrates clinical symptoms, cognitive performance, and brain structure measures to subtype Attention-Deficit/Hyperactivity Disorder (ADHD) and Obsessive-Compulsive Disorder (OCD). Our method leverages Polya-Gamma augmentation for computational efficiency and robust inference. Simulations demonstrate improved stability and accuracy compared to existing clustering approaches. Application to the ABCD data reveals clinically interpretable subtypes of ADHD and OCD with distinct cognitive and neurodevelopmental profiles. These findings show how integrative multimodal modeling can enhance the reproducibility and clinical relevance of psychiatric subtyping, supporting data-driven policies for early identification and targeted interventions in mental health.

stat.ME

Joint modeling for learning decision-making dynamics in behavioral experiments

Major depressive disorder (MDD), a leading cause of disability and mortality, is associated with reward-processing abnormalities and concentration issues. Motivated by the probabilistic reward task from the Establishing Moderators and Biosignatures of Antidepressant Response in Clinical Care (EMBARC) study, we propose a novel framework that integrates the reinforcement learning (RL) model and drift-diffusion model (DDM) to jointly analyze reward-based decision-making with response times. To account for emerging evidence suggesting that decision-making may alternate between multiple interleaved strategies, we model latent state switching using a hidden Markov model (HMM). In the ''engaged'' state, decisions follow an RL-DDM, simultaneously capturing reward processing, decision dynamics, and temporal structure. In contrast, in the ''lapsed'' state, decision-making is modeled using a simplified DDM, where specific parameters are fixed to approximate random guessing with equal probability. The proposed method is implemented using a computationally efficient generalized expectation-maximization (EM) algorithm with forward-backward procedures. Through extensive numerical studies, we demonstrate that our proposed method outperforms competing approaches across various reward-generating distributions, under both strategy-switching and non-switching scenarios, as well as in the presence of input perturbations. When applied to the EMBARC study, our framework reveals that MDD patients exhibit lower overall engagement than healthy controls and experience longer decision times when they do engage. Additionally, we show that neuroimaging measures of brain activities are associated with decision-making characteristics in the ''engaged'' state but not in the ''lapsed'' state, providing evidence of brain-behavior association specific to the ''engaged'' state.

stat.ME

Dynamic Classification of Latent Disease Progression with Auxiliary Surrogate Labels

Disease progression prediction based on patients' evolving health information is challenging when true disease states are unknown due to diagnostic capabilities or high costs. For example, the absence of gold-standard neurological diagnoses hinders distinguishing Alzheimer's disease (AD) from related conditions such as AD-related dementias (ADRDs), including Lewy body dementia (LBD). Combining temporally dependent surrogate labels and health markers may improve disease prediction. However, existing literature models informative surrogate labels and observed variables that reflect the underlying states using purely generative approaches, limiting the ability to predict future states. We propose integrating the conventional hidden Markov model as a generative model with a time-varying discriminative classification model to simultaneously handle potentially misspecified surrogate labels and incorporate important markers of disease progression. We develop an adaptive forward-backward algorithm with subjective labels for estimation, and utilize the modified posterior and Viterbi algorithms to predict the progression of future states or new patients based on objective markers only. Importantly, the adaptation eliminates the need to model the marginal distribution of longitudinal markers, a requirement in traditional algorithms. Asymptotic properties are established, and significant improvement with finite samples is demonstrated via simulation studies. Analysis of the neuropathological dataset of the National Alzheimer's Coordinating Center (NACC) shows much improved accuracy in distinguishing LBD from AD.

stat.ME

HMM for Discovering Decision-Making Dynamics Using Reinforcement Learning Experiments

Major depressive disorder (MDD) presents challenges in diagnosis and treatment due to its complex and heterogeneous nature. Emerging evidence indicates that reward processing abnormalities may serve as a behavioral marker for MDD. To measure reward processing, patients perform computer-based behavioral tasks that involve making choices or responding to stimulants that are associated with different outcomes. Reinforcement learning (RL) models are fitted to extract parameters that measure various aspects of reward processing to characterize how patients make decisions in behavioral tasks. Recent findings suggest the inadequacy of characterizing reward learning solely based on a single RL model; instead, there may be a switching of decision-making processes between multiple strategies. An important scientific question is how the dynamics of learning strategies in decision-making affect the reward learning ability of individuals with MDD. Motivated by the probabilistic reward task (PRT) within the EMBARC study, we propose a novel RL-HMM framework for analyzing reward-based decision-making. Our model accommodates learning strategy switching between two distinct approaches under a hidden Markov model (HMM): subjects making decisions based on the RL model or opting for random choices. We account for continuous RL state space and allow time-varying transition probabilities in the HMM. We introduce a computationally efficient EM algorithm for parameter estimation and employ a nonparametric bootstrap for inference. We apply our approach to the EMBARC study to show that MDD patients are less engaged in RL compared to the healthy controls, and engagement is associated with brain activities in the negative affect circuitry during an emotional conflict task.

cs.LG

Learning Optimal Dynamic Treatment Regimens Subject to Stagewise Risk Controls

Dynamic treatment regimens (DTRs) aim at tailoring individualized sequential treatment rules that maximize cumulative beneficial outcomes by accommodating patients' heterogeneity in decision-making. For many chronic diseases including type 2 diabetes mellitus (T2D), treatments are usually multifaceted in the sense that aggressive treatments with a higher expected reward are also likely to elevate the risk of acute adverse events. In this paper, we propose a new weighted learning framework, namely benefit-risk dynamic treatment regimens (BR-DTRs), to address the benefit-risk trade-off. The new framework relies on a backward learning procedure by restricting the induced risk of the treatment rule to be no larger than a pre-specified risk constraint at each treatment stage. Computationally, the estimated treatment rule solves a weighted support vector machine problem with a modified smooth constraint. Theoretically, we show that the proposed DTRs are Fisher consistent, and we further obtain the convergence rates for both the value and risk functions. Finally, the performance of the proposed method is demonstrated via extensive simulation studies and application to a real study for T2D patients.

stat.ME

Fusing Individualized Treatment Rules Using Secondary Outcomes

An individualized treatment rule (ITR) is a decision rule that recommends treatments for patients based on their individual feature variables. In many practices, the ideal ITR for the primary outcome is also expected to cause minimal harm to other secondary outcomes. Therefore, our objective is to learn an ITR that not only maximizes the value function for the primary outcome, but also approximates the optimal rule for the secondary outcomes as closely as possible. To achieve this goal, we introduce a fusion penalty to encourage the ITRs based on different outcomes to yield similar recommendations. Two algorithms are proposed to estimate the ITR using surrogate loss functions. We prove that the agreement rate between the estimated ITR of the primary outcome and the optimal ITRs of the secondary outcomes converges to the true agreement rate faster than if the secondary outcomes are not taken into consideration. Furthermore, we derive the non-asymptotic properties of the value function and misclassification rate for the proposed method. Finally, simulation studies and a real data example are used to demonstrate the finite-sample performance of the proposed method.

stat.ME

A Hierarchical Random Effects State-space Model for Modeling Brain Activities from Electroencephalogram Data

Mental disorders present challenges in diagnosis and treatment due to their complex and heterogeneous nature. Electroencephalogram (EEG) has shown promise as a potential biomarker for these disorders. However, existing methods for analyzing EEG signals have limitations in addressing heterogeneity and capturing complex brain activity patterns between regions. This paper proposes a novel random effects state-space model (RESSM) for analyzing large-scale multi-channel resting-state EEG signals, accounting for the heterogeneity of brain connectivities between groups and individual subjects. We incorporate multi-level random effects for temporal dynamical and spatial mapping matrices and address nonstationarity so that the brain connectivity patterns can vary over time. The model is fitted under a Bayesian hierarchical model framework coupled with a Gibbs sampler. Compared to previous mixed-effects state-space models, we directly model high-dimensional random effects matrices without structural constraints and tackle the challenge of identifiability. Through extensive simulation studies, we demonstrate that our approach yields valid estimation and inference. We apply RESSM to a multi-site clinical trial of Major Depressive Disorder (MDD). Our analysis uncovers significant differences in resting-state brain temporal dynamics among MDD patients compared to healthy individuals. In addition, we show the subject-level EEG features derived from RESSM exhibit a superior predictive value for the heterogeneous treatment effect compared to the EEG frequency band power, suggesting the potential of EEG as a valuable biomarker for MDD.

stat.ME

Learning Optimal Biomarker-Guided Treatment Policy for Chronic Disorders

Electroencephalogram (EEG) provides noninvasive measures of brain activity and is found to be valuable for diagnosis of some chronic disorders. Specifically, pre-treatment EEG signals in alpha and theta frequency bands have demonstrated some association with anti-depressant response, which is well-known to have low response rate. We aim to design an integrated pipeline that improves the response rate of major depressive disorder patients by developing an individualized treatment policy guided by the resting state pre-treatment EEG recordings and other treatment effects modifiers. We first design an innovative automatic site-specific EEG preprocessing pipeline to extract features that possess stronger signals compared with raw data. We then estimate the conditional average treatment effect using causal forests, and use a doubly robust technique to improve the efficiency in the estimation of the average treatment effect. We present evidence of heterogeneity in the treatment effect and the modifying power of EEG features as well as a significant average treatment effect, a result that cannot be obtained by conventional methods. Finally, we employ an efficient policy learning algorithm to learn an optimal depth-2 treatment assignment decision tree and compare its performance with Q-Learning and outcome-weighted learning via simulation studies and an application to a large multi-site, double-blind randomized controlled clinical trial, EMBARC.

stat.AP

Mixed-Response State-Space Model for Analyzing Multi-Dimensional Digital Phenotypes

Digital technologies (e.g., mobile phones) can be used to obtain objective, frequent, and real-world digital phenotypes from individuals. However, modeling these data poses substantial challenges since observational data are subject to confounding and various sources of variabilities. For example, signals on patients' underlying health status and treatment effects are mixed with variation due to the living environment and measurement noises. The digital phenotype data thus shows extensive variabilities between- and within-patient as well as across different health domains (e.g., motor, cognitive, and speaking). Motivated by a mobile health study of Parkinson's disease (PD), we develop a mixed-response state-space (MRSS) model to jointly capture multi-dimensional, multi-modal digital phenotypes and their measurement processes by a finite number of latent state time series. These latent states reflect the dynamic health status and personalized time-varying treatment effects and can be used to adjust for informative measurements. For computation, we use the Kalman filter for Gaussian phenotypes and importance sampling with Laplace approximation for non-Gaussian phenotypes. We conduct comprehensive simulation studies and demonstrate the advantage of MRSS in modeling a mobile health study that remotely collects real-time digital phenotypes from PD patients.

stat.AP

Learning non-monotone optimal individualized treatment regimes

We propose a new modeling and estimation approach to select the optimal treatment regime from different options through constructing a robust estimating equation. The method is protected against misspecification of the propensity score model, the outcome regression model for the non-treated group, or the potential non-monotonic treatment difference model. Our method also allows residual errors to depend on covariates. A single index structure is incorporated to facilitate the nonparametric estimation of the treatment difference. We then identify the optimal treatment through maximizing the value function. Theoretical properties of the treatment assignment strategy are established. We illustrate the performance and effectiveness of our proposed estimators through extensive simulation studies and a real dataset on the effect of maternal smoking on baby birth weight.

stat.ME