SearcharxivSearch

arXiv subjects

Zeyi Wang

Publications and source records attributed to Zeyi Wang.

10 recordsLinked to original sources

Higher-Order Efficient Estimators: A Review and Simulation-Based Benchmark Study

Higher-order efficient estimators extend standard first-order semiparametric estimators by replacing second-order residuals with third- or higher-order terms, potentially enabling asymptotic efficiency under slower nuisance function convergence rates and improving finite-sample performance. Existing methods achieve higher-order expansions through structurally different approximation strategies, including basis truncation, kernel smoothing, and highly adaptive lasso (HAL) representations, making direct theoretical and practical comparison difficult. In this manuscript, we provide a focused review and a simulation-based empirical benchmark for second-order efficient estimators, using treatment-specific mean estimation as a canonical causal inference and missing data problem. We compare how higher-order influence function (HOIF) estimators, kernel-based higher-order targeted minimum loss-based estimator (HOTMLE), and HAL-based HOTMLE construct higher-order expansions and the approximation or regularization burdens they introduce. The asymptotic and numerical study evaluates first-order and empirical second-order estimators under controlled nuisance errors with constant or increasing sectional variation complexity. Results show that higher-order debiasing can substantially reduce first-order estimation bias; however, gains depend strongly on stability of the approximation or regularization required for higher-order correction. Empirical HAL-based HOTMLE shows relatively stable performance, while empirical HOIF remains sensitive to basis truncation and tuning choices. Overall, this manuscript clarifies when higher-order asymptotic improvements are attained in theory, when they may be practically visible, and when implementation instability may offset theoretical advantages.

stat.ME

SN2024abfl: A Low-Luminosity Type IIP Supernova at the Low-Mass End of Core Collapse

We present optical photometric and spectroscopic observations of the low-luminosity (LL) Type IIP supernova SN\,2024abfl. The distance to its host galaxy is highly uncertain, with independent estimates of $9.5^{+2.3}_{-2.4}$ Mpc and $15.0^{+8.9}_{-1.9}$ Mpc. Even adopting the larger distance, the inferred plateau luminosity is only $\sim 10^{41}\rm erg\,s^{-1}$, placing SN 2024abfl at the extreme faint end of SNe IIP population. Its light curve exhibits a long-lasting plateau of approximately 110 days. The spectra show exceptionally low expansion velocities, with the \FeII\, velocity of $\sim1200\,\rm km\,s^{-1}$ at 50 days after the explosion, significantly lower than the typical values of $\sim2000-5500\,\rm km\,s^{-1}$ observed in SNe IIP, placing SN\,2024abfl among the slowest-expanding LL SNe IIP. Bolometric modeling yields a synthesized $^{56}$Ni mass of $\sim0.002-0.004\,\rm M_\odot$, though this estimate remains subject to significant uncertainty owing to the poorly constrained distance. Considering the plateau color and duration, the magnitude drop from plateau to tail, and the progenitor luminosity, we favor a low-mass core-collapse origin for SN\,2024abfl.

astro-ph.SR

SN 2023axu: A Type IIP Supernova Interacted with a Low-Density Stellar Wind

We present photometric and spectroscopic observations of Type IIP supernova SN 2023axu, spanning $\sim$400 d after the explosion. Its light curve is typical of normal SNe IIP, with a V-band peak of $-17.25 \pm 0.06$ mag and no early-time excess indicative of strong circumstellar interaction. The early spectra exhibit a distinctive broad "ledge" near 4600 Å. Through spectral modeling and comparison, we attribute this feature to a blend of C, N, and He lines excited by weak interaction between the ejecta and a low-density stellar wind. The late-time photometric evolution shows no discernible contribution from interaction, arguing against strong late-time circumstellar material engagement and supporting the low-density wind scenario. From modeling, this SN synthesized $\sim 0.055\,M_\odot$ of $^{56}$Ni, and nebular spectrum analysis indicates a progenitor mass near $15\,M_\odot$. SN 2023axu thus exemplifies weak ejecta-wind interaction and highlights the diversity of mass-loss histories and circumstellar environments of SNe II progenitors.

astro-ph.HE

Super Ensemble Learning Using the Highly-Adaptive-Lasso

We introduce the Meta Highly-Adaptive-Lasso Minimum Loss Estimator (M-HAL-MLE), a novel ensemble approach for estimating functional parameters of realistically modeled data distribution from independent and identically distributed observations. Given $J$ initial estimators, candidate ensembles are generated by finite-sectional-variation cadlag functions. Using $V$-fold cross-validation, the M-HAL-MLE selects the optimal cadlag ensemble minimizing the cross-validated empirical risk, with the sectional variation bound as a tuning parameter. The final estimator, M-HAL super-learner, is obtained by averaging ensemble compositions across folds. In contrast, the oracle ensemble and oracle estimator are defined by minimizing the population excess risk relative to the true function. We establish following theoretical properties: 1) the M-HAL super-learner converges to the oracle estimator at rate $n^{-2/3}$ in excess risk, up to log-n factors; 2) by appropriate undersmoothing, target features of the M-HAL super-learner are asymptotically linear for corresponding target features of the oracle estimator; 3) the excess risk between the oracle estimator and true function, along with the difference between their target features, is generally second-order. Simulations validate the theoretical results, demonstrating effectiveness in high-dimensional settings. We further illustrate the method in a real-data application involving mediation analysis of functional MRI from human pain studies.

stat.ME

Regularized Targeted Maximum Likelihood Estimation in Highly Adaptive Lasso Implied Working Models

We address the challenge of performing Targeted Maximum Likelihood Estimation (TMLE) after an initial Highly Adaptive Lasso (HAL) fit. Existing approaches that utilize the data-adaptive working model selected by HAL-such as the relaxed HAL update-can be simple and versatile but may become computationally unstable when the HAL basis expansions introduce collinearity. Undersmoothed HAL may fail to solve the efficient influence curve (EIC) at the desired level without overfitting, particularly in complex settings like survival-curve estimation. A full HAL-TMLE, which treats HAL as the initial estimator and then targets in the nonparametric or semiparametric model, typically demands costly iterative clever-covariate calculations in complex set-ups like survival analysis and longitudinal mediation analysis. To overcome these limitations, we propose two new HAL-TMLEs that operate within the finite-dimensional working model implied by HAL: Delta-method regHAL-TMLE and Projection-based regHAL-TMLE. We conduct extensive simulations to demonstrate the performance of our proposed methods.

stat.ME

Statistical Analysis of Data Repeatability Measures

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the characteristics of the data which are repeatable -- the aspects of the data that are able to be identified under a duplicated analysis. Conflictingly, the utility of traditional repeatability measures, such as the intraclass correlation coefficient, under these settings is limited. In recent work, novel data repeatability measures have been introduced in the context where a set of subjects are measured twice or more, including: fingerprinting, rank sums, and generalizations of the intraclass correlation coefficient. However, the relationships between, and the best practices among these measures remains largely unknown. In this manuscript, we formalize a novel repeatability measure, discriminability. We show that it is deterministically linked with the correlation coefficient under univariate random effect models, and has desired property of optimal accuracy for inferential tasks using multivariate measurements. Additionally, we overview and systematically compare repeatability statistics using both theoretical results and simulations. We show that the rank sum statistic is deterministically linked to a consistent estimator of discriminability. The power of permutation tests derived from these measures are compared numerically under Gaussian and non-Gaussian settings, with and without simulated batch effects. Motivated by both theoretical and empirical results, we provide methodological recommendations for each benchmark setting to serve as a resource for future analyses. We believe these recommendations will play an important role towards improving repeatability in fields such as functional magnetic resonance imaging, genomics, pharmacology, and more.

stat.AP

Large Language Model-Aided Evolutionary Search for Constrained Multiobjective Optimization

Evolutionary algorithms excel in solving complex optimization problems, especially those with multiple objectives. However, their stochastic nature can sometimes hinder rapid convergence to the global optima, particularly in scenarios involving constraints. In this study, we employ a large language model (LLM) to enhance evolutionary search for solving constrained multi-objective optimization problems. Our aim is to speed up the convergence of the evolutionary population. To achieve this, we finetune the LLM through tailored prompt engineering, integrating information concerning both objective values and constraint violations of solutions. This process enables the LLM to grasp the relationship between well-performing and poorly performing solutions based on the provided input data. Solution's quality is assessed based on their constraint violations and objective-based performance. By leveraging the refined LLM, it can be used as a search operator to generate superior-quality solutions. Experimental evaluations across various test benchmarks illustrate that LLM-aided evolutionary search can significantly accelerate the population's convergence speed and stands out competitively against cutting-edge evolutionary algorithms.

cs.NE

Applying the causal roadmap to longitudinal national Danish registry data: a case study of second-line diabetes medication and dementia

The causal roadmap is a formal framework for causal and statistical inference that supports clear specification of the causal question, interpretable and transparent statement of required causal assumptions, robust inference, and optimal precision. The roadmap is thus particularly well-suited to evaluating longitudinal causal effects using large scale registries; however, application of the roadmap to registry data also introduces particular challenges. In this paper we provide a detailed case study of the longitudinal causal roadmap applied to the Danish National Registry to evaluate the comparative effectiveness of second-line diabetes drugs on dementia risk. Specifically, we evaluate the difference in counterfactual five-year cumulative risk of dementia if a target population of adults with type 2 diabetes had initiated and remained on GLP-1 receptor agonists (a second-line diabetes drug) compared to a range of active comparator protocols. Time-dependent confounding is accounted for through use of the iterated conditional expectation representation of the longitudinal g-formula as a statistical estimand. Statistical estimation uses longitudinal targeted maximum likelihood, incorporating machine learning. We provide practical guidance on the implementation of the roadmap using registry data, and highlight how rare exposures and outcomes over long-term follow up can raise challenges for flexible and robust estimators, even in the context of the large sample sizes provided by the registry. We demonstrate how simulations can be used to help address these challenges by supporting careful estimator pre-specification. We find a protective effect of GLP-1RAs compared to some but not all other second-line treatments.

stat.AP

Targeted Maximum Likelihood Based Estimation for Longitudinal Mediation Analysis

Causal mediation analysis with random interventions has become an area of significant interest for understanding time-varying effects with longitudinal and survival outcomes. To tackle causal and statistical challenges due to the complex longitudinal data structure with time-varying confounders, competing risks, and informative censoring, there exists a general desire to combine machine learning techniques and semiparametric theory. In this manuscript, we focus on targeted maximum likelihood estimation (TMLE) of longitudinal natural direct and indirect effects defined with random interventions. The proposed estimators are multiply robust, locally efficient, and directly estimate and update the conditional densities that factorize data likelihoods. We utilize the highly adaptive lasso (HAL) and projection representations to derive new estimators (HAL-EIC) of the efficient influence curves of longitudinal mediation problems and propose a fast one-step TMLE algorithm using HAL-EIC while preserving the asymptotic properties. The proposed method can be generalized for other longitudinal causal parameters that are smooth functions of data likelihoods, and thereby provides a novel and flexible statistical toolbox.

stat.ME

Higher Order Targeted Maximum Likelihood Estimation

Asymptotic efficiency of targeted maximum likelihood estimators (TMLE) of target features of the data distribution relies on a a second order remainder being asymptotically negligible. In previous work we proposed a nonparametric MLE termed Highly Adaptive Lasso (HAL) which parametrizes the relevant functional of the data distribution in terms of a multivariate real valued cadlag function that is assumed to have finite variation norm. We showed that the HAL-MLE converges in Kullback-Leibler dissimilarity at a rate n-1/3 up till logn factors. Therefore, by using HAL as initial density estimator in the TMLE, the resulting HAL-TMLE is an asymptotically efficient estimator only assuming that the relevant nuisance functions of the data density are cadlag and have finite variation norm. However, in finite samples, the second order remainder can dominate the sampling distribution so that inference based on asymptotic normality would be anti-conservative. In this article we propose a new higher order TMLE, generalizing the regular first order TMLE. We prove that it satisfies an exact linear expansion, in terms of efficient influence functions of sequentially defined higher order fluctuations of the target parameter, with a remainder that is a k+1th order remainder. As a consequence, this k-th order TMLE allows statistical inference only relying on the k+1th order remainder being negligible. We also provide a rationale for the higher order TMLE that it will be superior to the first order TMLE by (iteratively) locally minimizing the exact finite sample remainder of the first order TMLE. The second order TMLE is demonstrated for nonparametric estimation of the integrated squared density and for the treatment specific mean outcome. We also provide an initial simulation study for the second order TMLE of the treatment specific mean confirming the theoretical analysis.

math.ST