SearcharxivSearch

arXiv subjects

Jianbin Tan

Publications and source records attributed to Jianbin Tan.

11 recordsLinked to original sources

Panel Flow Matching: A Generative Approach to Learning Distributions of Longitudinal Data

Learning distributions of longitudinal data is central to tasks such as visualization, completion, classification, and synthetic data generation, but it remains statistically challenging because longitudinal observations are often irregular, sparse, and collected from only a limited number of subjects. To address this, we develop a novel generative framework, termed panel flow matching (PFM), for learning longitudinal distributions by pooling information across time via a continuous panel flow model. PFM combines a forward flow-matching step with a backward kernel-fitting step, yielding a flexible and data-adaptive approach for capturing complex distributional structures. We apply PFM to estimate panel densities, namely the cross-sectional densities of longitudinal data, and establish statistical guarantees under irregular and sparse sampling designs. Under this, PFM naturally supports tasks including longitudinal completion, synthetic data generation, and classification, without requiring a preliminary dimension-reduction step to handle data irregularity. Extensive simulations demonstrate that PFM outperforms existing methods across these tasks. We further apply PFM to a vaginal microbiome longitudinal dataset from 188 pregnancies labeled as term or preterm, where it improves classification accuracy and reveals time-varying distributional differences between the two groups.

stat.ME

A Frequency-Domain Approach for Integrating Multiple Functional Time Series

Integrative analysis of multivariate functional time series (MFTS) is both critical and challenging across many scientific domains. Such data often exhibit complex multi-way dependencies arising from within-curve structures, temporal correlations across curves, and cross-subject interactions, underscoring the need for efficient methods that can jointly capture these dependencies and support accurate downstream analyses. In this work, we propose a novel frequency-domain framework based on a marginal dynamic Karhunen--Lo\`eve expansion. The key idea is to integrate individual spectral densities of the MFTS to construct a marginal spectral operator, whose eigenfunctions yield optimal functional filters. These filters transform complex functional observations into a structured multivariate time series representation, providing a powerful foundation for joint modeling and estimation. Through extensive simulation studies, we demonstrate the superior performance of the proposed approach. We further validate its practical utility through an application to the imputation and forecasting of air pollutant concentration trajectories in China.

stat.ME

Associating High-Dimensional Longitudinal Datasets through an Efficient Cross-Covariance Decomposition

Understanding associations between paired high-dimensional longitudinal datasets is a fundamental yet challenging problem that arises across scientific domains, including longitudinal multi-omic studies. The difficulty stems from the complex, time-varying cross-covariance structure coupled with high dimensionality, which complicates both model formulation and statistical estimation. To address these challenges, we propose a new framework, termed Functional-Aggregated Cross-covariance Decomposition (FACD), tailored for canonical cross-covariance analysis between paired high-dimensional longitudinal datasets through a statistically efficient and theoretically grounded procedure. Unlike existing methods that are often limited to low-dimensional data or rely on explicit parametric modeling of temporal dynamics, FACD adaptively learns temporal structure by aggregating signals across features and naturally accommodates variable selection to identify the most relevant features associated across datasets. We establish statistical guarantees for FACD and demonstrate its advantages over existing approaches through extensive simulation studies. Finally, we apply FACD to a longitudinal multi-omic human study, revealing blood molecules with time-varying associations across omic layers during acute exercise.

stat.ME

Smooth Flow Matching for Synthesizing Functional Data

Functional data, i.e., random functions observed over a continuous domain, are increasingly available in areas such as biomedical research, health informatics, and epidemiology. However, effective statistical analysis for functional data is often hindered by challenges such as privacy constraints, sparse and irregular sampling, infinite-dimensionality, and non-Gaussian structures. To address these challenges, we introduce a novel framework named Smooth Flow Matching (SFM), tailored for generative modeling of functional data that enables statistical analysis without exposing sensitive real data. Under a copula framework, SFM constructs a semiparametric smooth flow to generate infinite-dimensional functional data, free of Gaussianity and low-rank assumptions. It is computationally efficient, handles irregular observations, and guarantees the smoothness of the generated functions, offering a practical and flexible solution in scenarios where existing deep generative methods are not applicable. Through extensive simulation studies, we demonstrate the advantages of SFM in terms of both synthetic data quality and computational efficiency. We then apply SFM to generate clinical trajectory data from the MIMIC-IV patient electronic health records (EHR) longitudinal database. Our analysis showcases the ability of SFM to produce high-quality surrogate data for downstream tasks, highlighting its potential to boost the utility of EHR data for clinical applications.

stat.ML

Sparse Equation Matching: A Derivative-Free Learning for General-Order Dynamical Systems

Equation discovery is a fundamental learning task for uncovering the underlying dynamics of complex systems, with wide-ranging applications in areas such as brain connectivity analysis, climate modeling, gene regulation, and physical simulation. However, many existing approaches rely on accurate derivative estimation and are limited to first-order dynamical systems, restricting their applicability in real-world scenarios. In this work, we propose Sparse Equation Matching (SEM), a unified framework that encompasses several existing equation discovery methods under a common formulation. SEM introduces an integral-based sparse regression approach using Green's functions, enabling derivative-free estimation of differential operators and their associated driving functions in general-order dynamical systems. The effectiveness of SEM is demonstrated through extensive simulations, benchmarking its performance against derivative-based approaches. We then apply SEM to electroencephalographic (EEG) data recorded during multiple oculomotor tasks, collected from 52 participants in a brain-computer interface experiment. Our method identifies active brain regions across participants and reveals task-specific connectivity patterns. These findings offer valuable insights into brain connectivity and the underlying neural mechanisms.

cs.LG

Integrated Analysis for Electronic Health Records with Structured and Sporadic Missingness

Objectives: We propose a novel imputation method tailored for Electronic Health Records (EHRs) with structured and sporadic missingness. Such missingness frequently arises in the integration of heterogeneous EHR datasets for downstream clinical applications. By addressing these gaps, our method provides a practical solution for integrated analysis, enhancing data utility and advancing the understanding of population health. Materials and Methods: We begin by demonstrating structured and sporadic missing mechanisms in the integrated analysis of EHR data. Following this, we introduce a novel imputation framework, Macomss, specifically designed to handle structurally and heterogeneously occurring missing data. We establish theoretical guarantees for Macomss, ensuring its robustness in preserving the integrity and reliability of integrated analyses. To assess its empirical performance, we conduct extensive simulation studies that replicate the complex missingness patterns observed in real-world EHR systems, complemented by validation using EHR datasets from the Duke University Health System (DUHS). Results: Simulation studies show that our approach consistently outperforms existing imputation methods. Using datasets from three hospitals within DUHS, Macomss achieves the lowest imputation errors for missing data in most cases and provides superior or comparable downstream prediction performance compared to benchmark methods. Conclusions: We provide a theoretically guaranteed and practically meaningful method for imputing structured and sporadic missing data, enabling accurate and reliable integrated analysis across multiple EHR datasets. The proposed approach holds significant potential for advancing research in population health.

stat.AP

Functional-SVD for Heterogeneous Trajectories: Case Studies in Health

Trajectory data, including time series and longitudinal measurements, are increasingly common in health-related domains such as biomedical research and epidemiology. Real-world trajectory data frequently exhibit heterogeneity across subjects such as patients, sites, and subpopulations, yet many traditional methods are not designed to accommodate such heterogeneity in data analysis. To address this, we propose a unified framework, termed Functional Singular Value Decomposition (FSVD), for statistical learning with heterogeneous trajectories. We establish the theoretical foundations of FSVD and develop a corresponding estimation algorithm that accommodates noisy and irregular observations. We further adapt FSVD to a wide range of trajectory-learning tasks, including dimension reduction, factor modeling, regression, clustering, and data completion, while preserving its ability to account for heterogeneity, leverage inherent smoothness, and handle irregular sampling. Through extensive simulations, we demonstrate that FSVD-based methods consistently outperform existing approaches across these tasks. Finally, we apply FSVD to a COVID-19 case-count dataset and electronic health record datasets, showcasing its effective performance in global and subgroup pattern discovery and factor analysis.

stat.ME

A Unified Principal Components Analysis for Stationary Functional Time Series

Functional time series (FTS) data have become increasingly available in real-world applications. Research on such data typically focuses on two objectives: curve reconstruction and forecasting, both of which require efficient dimension reduction. While functional principal component analysis (FPCA) serves as a standard tool, existing methods often fail to achieve simultaneous parsimony and optimality in dimension reduction, thereby restricting their practical implementation. To address this limitation, we propose a novel notion termed optimal functional filters, which unifies and enhances conventional FPCA methodologies. Specifically, we establish connections among diverse FPCA approaches through a dependence-adaptive representer for stationary FTS. Building on this theoretical foundation, we develop an estimation procedure for optimal functional filters that enables both dimension reduction and prediction within a Bayesian modeling framework. Theoretical properties are established for the proposed methodology, and comprehensive simulation studies validate its superiority over competing approaches. We further illustrate our method through an application to reconstructing and forecasting daily air pollutant concentration trajectories.

stat.ME

Functional Clustering for Longitudinal Associations between Social Determinants of Health and Stroke Mortality in the US

Understanding the longitudinally changing associations between Social Determinants of Health (SDOH) and stroke mortality is essential for effective stroke management. Previous studies have uncovered significant regional disparities in the relationships between SDOH and stroke mortality. However, existing studies have not utilized longitudinal associations to develop data-driven methods for regional division in stroke control. To fill this gap, we propose a novel clustering method to analyze SDOH -- stroke mortality associations in US counties. To enhance the interpretability of the clustering outcomes, we introduce a novel regularized expectation-maximization algorithm equipped with various sparsity-and-smoothness-pursued penalties, aiming at simultaneous clustering and variable selection in longitudinal associations. As a result, we can identify crucial SDOH that contribute to longitudinal changes in stroke mortality. This facilitates the clustering of US counties into different regions based on the relationships between these SDOH and stroke mortality. The effectiveness of our proposed method is demonstrated through extensive numerical studies. By applying our method to longitudinal data on SDOH and stroke mortality at the county level, we identify 18 important SDOH for stroke mortality and divide the US counties into two clusters based on these selected SDOH. Our findings unveil complex regional heterogeneity in the longitudinal associations between SDOH and stroke mortality, providing valuable insights into region-specific SDOH adjustments for mitigating stroke mortality.

stat.ME

Green's matching: an efficient approach to parameter estimation in complex dynamic systems

Parameters of differential equations are essential to characterize intrinsic behaviors of dynamic systems. Numerous methods for estimating parameters in dynamic systems are computationally and/or statistically inadequate, especially for complex systems with general-order differential operators, such as motion dynamics. This article presents Green's matching, a computationally tractable and statistically efficient two-step method, which only needs to approximate trajectories in dynamic systems but not their derivatives due to the inverse of differential operators by Green's function. This yields a statistically optimal guarantee for parameter estimation in general-order equations, a feature not shared by existing methods, and provides an efficient framework for broad statistical inferences in complex dynamic systems.

stat.ME

Graphical Principal Component Analysis of Multivariate Functional Time Series

In this paper, we consider multivariate functional time series with a two-way dependence structure: a serial dependence across time points and a graphical interaction among the multiple functions within each time point. We develop the notion of dynamic weak separability, a more general condition than those assumed in literature, and use it to characterize the two-way structure in multivariate functional time series. Based on the proposed weak separability, we develop a unified framework for functional graphical models and dynamic principal component analysis, and further extend it to optimally reconstruct signals from contaminated functional data using graphical-level information. We investigate asymptotic properties of the resulting estimators and illustrate the effectiveness of our proposed approach through extensive simulations. We apply our method to hourly air pollution data that were collected from a monitoring network in China.

stat.ME