SearcharxivSearch

arXiv subjects

Li-Hsiang Lin

Publications and source records attributed to Li-Hsiang Lin.

6 recordsLinked to original sources

AdaptICA: Data-Adaptive Transformation Learning for Independent Component Analysis

Independent component analysis (ICA) is widely used to recover latent structure from signal and imaging data, but standard ICA assumes that the observed measurement scale preserves a linear mixing structure. This assumption may fail for features produced through nonlinear preprocessing, such as band-specific power in motor-imagery EEG. We propose AdaptICA, an adaptive transformation-based framework that jointly learns grouped componentwise transformations and the demixing structure using a profiled mutual-information criterion. Because the transformation and demixing parameters may compensate for one another, their joint estimation introduces new identifiability and asymptotic challenges. We establish identifiability, consistency, and asymptotic normality of the transformation estimator, together with joint strong consistency of the transformation and demixing estimators. AdaptICA selects the transformation structure data-adaptively and includes the identity transformation as a candidate, thereby reducing to standard ICA when no scale adjustment is needed. Extensive simulations support the theoretical results. Applications demonstrate that AdaptICA can recover more independent and interpretable sources when transformation is beneficial while retaining standard ICA when the original measurement scale is adequate.

stat.ME

Similarity-Based Prediction for Digital Twins: Panel Data, Theory, and Applications

Prediction from sequential panel data is central to digital-twin modeling, where new panels arrive over time and the predictive system is updated sequentially. Existing methods often rely on temporal proximity, which can fail when similar input-output patterns recur at nonadjacent times or when recent panels differ from the target panel. We propose State-Local Prediction (StaLoP), a nonparametric dynamic panel prediction framework that utilizes information through target-local predictive compatibility. StaLoP represents panels using target-local state vectors, compares historical and target panels via empirical discrepancy scores to determine relevance weights for the target point, and combines these weights with covariate localization. Theoretical results, including bias-variance characterization, asymptotic normality, simultaneous prediction bands, and a target-local-GDF-corrected MSPE criterion for panel and model selection, are developed. Extensive simulations validate the performance of StaLoP and support its theoretical properties. Applications to sequence prediction, simulator calibration, variable selection, and county-to-county migration-flow forecasting demonstrate improved out-of-sample prediction and provide scientific insights into the underlying applications.

stat.ME

Scalable and Communication-Efficient Varying Coefficient Mixed Effect Models: Methodology, Theory, and Applications

Human migration exhibits complex spatiotemporal dependence driven by environmental and socioeconomic forces. Modeling such patterns at scale requires methods that accommodate many random effects while remaining feasible when raw data or large design matrices cannot be freely shared across distributed nodes. We develop a communication-efficient inference framework for Varying Coefficient Mixed Models (VCMMs) with flexible mean structures and large correlated random-effect components. Using a Bayesian hierarchical representation of penalized splines, we derive sufficient statistics that preserve each node's likelihood contribution and recover the estimator from the full data under unrestricted communication. Under communication constraints, these statistics support a one-step communication-efficient estimator with first-order efficiency. An SVD-enhanced implementation stabilizes large or ill-conditioned random-effect covariance operators. Theory establishes likelihood preservation, convergence, asymptotic efficiency, and finite-sample concentration. Simulations and U.S. migration-flow data demonstrate accuracy, scalability, and recovery of dynamic spatial patterns.

stat.ME

Sparse Deep Additive Model with Interactions: Enhancing Interpretability and Predictability

Recent advances in deep learning highlight the need for personalized models that can learn from small samples, handle high-dimensional features, and remain interpretable. To address this, we propose the Sparse Deep Additive Model with Interactions (SDAMI), a framework that combines sparsity-driven feature selection with deep subnetworks for flexible function approximation. Central to SDAMI is the Effect Footprint principle, which posits that higher-order interactions leave detectable marginal traces on constituent variables, enabling their discovery without exhaustive search. SDAMI executes this principle through a three-stage strategy: (1) screening for footprint variables, (2) disentangling main effects from interactions via group lasso, and (3) modeling components with dedicated deep subnetworks. Theoretical analysis confirms that footprints vanish only under measure-zero symmetry conditions that are rare in practice, ensuring consistent interaction recovery. Extensive simulations demonstrate that SDAMI successfully identifies pure interactions that heredity-based baselines fundamentally miss, recovering complex effect structures with near-zero false positive rates. Together, these results position SDAMI as a principled framework for interpretable high-dimensional regression.

stat.ML

Deep P-Spline: Theory, Fast Tuning, and Application

Deep neural networks (DNNs) have been widely applied to solve real-world regression problems. However, selecting optimal network structures remains a significant challenge. This study addresses this issue by linking neuron selection in DNNs to knot placement in basis expansion techniques. We introduce a difference penalty that automates knot selection, thereby simplifying the complexities of neuron selection. We name this method Deep P-Spline (DPS). This approach extends the class of models considered in conventional DNN modeling and forms the basis for a latent variable modeling framework using the Expectation-Conditional Maximization (ECM) algorithm for efficient network structure tuning with theoretical guarantees. From a nonparametric regression perspective, DPS is proven to overcome the curse of dimensionality, enabling the effective handling of datasets with a large number of input variable, a scenario where conventional nonparametric regression methods typically underperform. This capability motivates the application of the proposed methodology to computer experiments and image data analyses, where the associated regression problems involving numerous inputs are common. Numerical results validate the effectiveness of the model, underscoring its potential for advanced nonlinear regression tasks.

stat.CO

Yakovlev Promotion Time Cure Model with Local Polynomial Estimation

In modeling survival data with a cure fraction, flexible modeling of covariate effects on the probability of cure has important medical implications, which aids investigators in identifying better treatments to cure. This paper studies a semiparametric form of the Yakovlev promotion time cure model that allows for nonlinear effects of a continuous covariate. We adopt the local polynomial approach and use the local likelihood criterion to derive nonlinear estimates of covariate effects on cure rates, assuming that the baseline distribution function follows a parametric form. This way we adopt a flexible method to estimate the cure rate locally, the important part in cure models, and a convenient way to estimate the baseline function globally. An algorithm is proposed to implement estimation at both the local and global scales. Asymptotic properties of local polynomial estimates, the nonparametric part, are investigated in the presence of both censored and cured data, and the parametric part is shown to be root-n consistent. The proposed methods are illustrated by simulated and real data.

stat.ME