SearcharxivSearch

arXiv subjects

Mengfei Ran

Publications and source records attributed to Mengfei Ran.

7 recordsLinked to original sources

Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional and temporal encoders map point clouds and prior histories into states; nuisance heads estimate outcome, treatment, and censoring functions; and EIF-targeted validation, calibration, overlap, tail, and ablation diagnostics assess whether the state supports the estimating equation. If the selected state preserves the nuisance information needed by the EIF, representation error enters the same second-order product remainder as ordinary nuisance error, and the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions. Catoni aggregation is treated separately as a bounded-influence point estimator, not a replacement for Wald inference. Simulations show gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed. A VitalDB audit shows that DR-FRL can use irregular laboratory point clouds and deliver a useful negative finding: for this ICU-disposition endpoint, scalar laboratory summaries already carry much endpoint-relevant information.

stat.ML

Semiparametric Inference for Causal Effects on Functional Outcomes

Difference-in-differences (DiD) is a cornerstone of causal inference, yet extending it to functional outcomes is not a routine scalar generalization; rather, it entails three fundamental challenges in identification, inference, and observation. This paper develops a comprehensive semiparametric inference framework for functional DiD with discretely observed data. First, we define the functional average treatment effect under parallel trends and derive its efficient influence function (EIF), thereby establishing the semiparametric efficiency bound. Second, leveraging Neyman orthogonality and cross-fitting, we construct a debiased estimator that effectively mitigates regularization bias arising from nonparametric reconstruction. Third, we establish weak convergence of the estimator and propose an asymptotically valid uniform confidence band, enabling a rigorous transition from pointwise to curve-level inference. Finally, we demonstrate that reconstruction error under discrete sampling is asymptotically negligible for semiparametric inference, ensuring practical feasibility. Simulations and empirical applications confirm that the proposed method achieves superior coverage and testing power in finite samples, providing a theoretically grounded and computationally tractable foundation for causal evaluation with functional data.

stat.ME

Group-Sparse Smoothing for Longitudinal Models with Time-Varying Coefficients

Longitudinal associations may vary over time, yet allowing every regression effect to be dynamic can inflate estimation variance and obscure interpretable structure. We develop time-varying-effect selection (TV-Select), a group-sparse smoothing framework that classifies covariate effects as zero, constant, or time varying. Each coefficient is decomposed into a constant mean and a centered temporal deviation represented by a full-rank, L2-normalized effective spline basis. A group penalty identifies varying components, while a roughness penalty controls their curvature. The resulting convex criterion is solved by cyclic block proximal-gradient updates and followed by smooth refitting. Under a full-column-rank unpenalized design and an effective model dimension that is small relative to the total number of observations, we establish prediction and parameter rates, blockwise function-estimation bounds, and exact recovery of the varying set under irrepresentability and beta-min conditions. A stable classification refit further separates zero from constant effects. For fixed-dimensional contrasts, we construct an oracle-equivalent one-step estimator with cluster-robust asymptotic normality and consistent sandwich variance estimation. Simulations demonstrate that TV-Select combines low false-positive rates with accurate function estimation and competitive prediction across a range of longitudinal settings. An application to Sleep-EDF data produces smooth and parsimonious temporal effect estimates with essentially unchanged held-out predictive performance.

stat.ME

Adaptive Penalized Doubly Robust Regression for Longitudinal Data

Longitudinal data often involve heterogeneity, sparse signals, and contamination from response outliers or high-leverage observations especially in biomedical science. Existing methods usually address only part of this problem, either emphasizing penalized mixed effects modeling without robustness or robust mixed effects estimation without high-dimensional variable selection. We propose a doubly adaptive robust regression (DAR-R) framework for longitudinal linear mixed effects models. It combines a robust pilot fit, doubly adaptive observation weights for residual outliers and leverage points, and folded concave penalization for fixed effect selection, together with weighted updates of random effects and variance components. We develop an iterative reweighting algorithm and establish estimation and prediction error bounds, support recovery consistency, and oracle-type asymptotic normality. Simulations show that DAR-R improves estimation accuracy, false-positive control, and covariance estimation under both vertical outliers and bad leverage contamination. In the TADPOLE/ADNI Alzheimer's disease application, DAR-R achieves accurate and stable prediction of ADAS13 while selecting clinically meaningful predictors with strong resampling stability.

stat.ME

Block Empirical Likelihood Inference for Longitudinal Generalized Partially Linear Single-Index Models

Generalized partially linear single-index models (GPLSIMs) provide a flexible and interpretable semiparametric framework for longitudinal outcomes by combining a low-dimensional parametric component with a nonparametric index component. For repeated measurements, valid inference is challenging because within-subject correlation induces nuisance parameters and variance estimation can be unstable in semiparametric settings. We propose a profile estimating-equation approach based on spline approximation of the unknown link function and construct a subject-level block empirical likelihood (BEL) for joint inference on the parametric coefficients and the single-index direction. The resulting BEL ratio statistic enjoys a Wilks-type chi-square limit, yielding likelihood-free confidence regions without explicit sandwich variance estimation. We also discuss practical implementation, including constrained optimization for the index parameter, working-correlation choices, and bootstrap-based confidence bands for the nonparametric component. Simulation studies and an application to the epilepsy longitudinal study illustrate the finite-sample performance.

stat.ME

A Generalized Adaptive Joint Learning Framework for High-Dimensional Time-Varying Models

In modern biomedical and econometric studies, longitudinal processes are often characterized by complex time-varying associations and abrupt regime shifts that are shared across correlated outcomes. Standard functional data analysis (FDA) methods, which prioritize smoothness, often fail to capture these dynamic structural features, particularly in high-dimensional settings. This article introduces Adaptive Joint Learning (AJL), a hierarchical regularization framework designed to integrate functional variable selection with structural changepoint detection in multivariate time-varying coefficient models. Unlike standard simultaneous estimation approaches, we propose a theoretically grounded two-stage screening-and-refinement procedure. This framework first synergizes adaptive group-wise penalization with sure screening principles to robustly identify active predictors, followed by a refined fused regularization step that effectively borrows strength across multiple outcomes to detect local regime shifts. We provide a rigorous theoretical analysis of the estimator in the ultra-high-dimensional regime (p >> n). Crucially, we establish the sure screening consistency of the first stage, which serves as the foundation for proving that the refined estimator achieves the oracle property-performing as well as if the true active set and changepoint locations were known a priori. A key theoretical contribution is the explicit handling of approximation bias via undersmoothing conditions to ensure valid asymptotic inference. The proposed method is validated through comprehensive simulations and an application to Sleep-EDF data, revealing novel dynamic patterns in physiological states.

stat.ME

Universal 2-Local Symmetry-Preserving Quantum Neural Networks for Fermionic Systems

Simulating quantum many-body systems represents a fundamental challenge where classical machine learning methods are severely bottlenecked by the exponential curse of dimensionality. Variational Quantum Algorithms (VQAs) offer a native paradigm to tackle this by optimizing parameterized unitary evolutions to find the ground states of problem Hamiltonians. However, the efficacy of these VQA is deeply hindered by the challenge of balancing the preservation of critical physical symmetries with the strict constraints of hardware implementability. In this work, we address this dilemma by proposing a hardware-efficient, symmetry-preserving ansatz fortified with complete theoretical guarantees for fermionic systems, termed the Hamming Weight Preserving (HWP) ansatz. We establish the necessary and sufficient conditions for 2-local HWP operators to achieve subspace universality, formally debunking the prevailing assumption that truncation-free simulation requires complex high-order interactions. Empirical validations corroborate our theoretical guarantees, showcasing the exact approximation of arbitrary unitary matrices within the HWP subspace. Crucially, we demonstrate the exceptional versatility of the proposed approach by deploying the exact same ansatz across distinct fermionic models, including diverse molecular electronic structures and the Fermi-Hubbard model. Our proposed HWP ansatz consistently suppresses ground-state energy errors below $1 \times 10^{-10}$ Ha, achieving a level of precision that surpasses the stringent threshold of chemical accuracy by multiple orders of magnitude. This work establishes a complete, theoretically fortified 2-local framework for symmetry-preserving computation, offering a highly universal and hardware-efficient building block for advancing quantum machine learning and fermionic many-body simulations.

quant-ph