SearcharxivSearch

arXiv subjects

Sihan Wu

Publications and source records attributed to Sihan Wu.

4 recordsLinked to original sources

Debiased inference for proximal dose-response function

In this paper, we study nonparametric inference for the causal dose-response curve of a continuous-treatment under unmeasured confounding by leveraging treatment- and outcome-inducing confounding proxies. To estimate the curve, we introduce a novel proximal doubly robust pseudo-outcome whose conditional mean given treatment equals the dose-response curve whenever either bridge function is correctly specified, thereby addressing a key gap in proximal causal inference for continuous-treatments. Furthermore, we derive an influence function for its smoothed causal estimand, and construct a cross-fitted debiased local-linear estimator with a proper local-quadratic bias correction. We establish pointwise and finite-dimensional asymptotic normality and a uniform Gaussian approximation over compact treatment intervals. Both smoothing bandwidths may have the mean-squared-error-optimal order without undersmoothing, while cross-fitting accommodates flexible bridge estimators under a product convergence rate conditions without fitted-class entropy restrictions. We also develop practical bandwidth selectors, pointwise confidence intervals, and simultaneous confidence bands. Extensive simulations and a data analysis highlight the practical performance of the proposed method under latent confounding and multiple proxies.

stat.ME

Proximal Mediation Analysis with Hidden Recanting Witnesses

Mediation analysis is essential for decomposing the causal effect of a treatment into direct and indirect pathways. However, many practical settings rely on the stringent assumption that recanting witnesses, defined as treatment-induced mediator-outcome confounders, are either absent or fully known a priori. Such a requirement is often untenable, especially when these variables remain unobservable due to measurement difficulties or privacy constraints. In this paper, we leverage proximal causal inference to develop three novel identification strategies to address the challenge of identifying path-specific effects in the presence of unknown recanting witnesses. Building on this, we develop a semiparametric inference framework that derives the efficient influence function and proposes a proximal multiply robust estimator, which remains consistent if at least one set of nuisance models is correctly specified. When all nuisance models are correctly specified and converge at appropriate rates, the estimator is asymptotically normal and achieves the semiparametric efficiency bound. We provide a minimax optimization-based debiased machine learning procedure for point estimation and constructing valid confidence intervals. The performance of the proposed methods is demonstrated by simulation studies and a real data application.

stat.ME

Proximal Path-Specific Inference

Causal mediation analysis has been extended to estimate path-specific effects with multiple intermediate variables, isolating treatment effects through a mediator of interest while excluding pathways through its ancestors. Such analyses address bias from recanting witnesses, i.e., treatment-induced mediator-outcome confounders. However, existing methods typically rely on stringent assumptions precluding general unmeasured confounding, which are often violated in practice. In this paper, we relax these restrictions by leveraging observed covariates as proxy variables to accommodate unmeasured confounding among the treatment, recanting witness, mediator, and outcome. Using proximal confounding bridge functions, we develop four nonparametric identification strategies for the path-specific effect. We further derive the efficient influence function and propose a quadruply robust, locally efficient estimator. To handle high-dimensional nuisance parameters, we propose a proximal debiased machine learning approach. We theoretically guarantee that our estimator achieves $\sqrt{n}$-consistency and asymptotic normality even when machine learning estimators for nuisance functions converge at slower rates. Our approaches are validated via semiparametric and nonparametric simulations and an application to the CDC WONDER Natality study, estimating the path-specific effect of prenatal care on preterm birth through preeclampsia, independent of maternal smoking during pregnancy.

stat.ME

Do Transformers Have the Ability for Periodicity Generalization?

Large language models (LLMs) based on the Transformer have demonstrated strong performance across diverse tasks. However, current models still exhibit substantial limitations in out-of-distribution (OOD) generalization compared with humans. We investigate this gap through periodicity, one of the basic OOD scenarios. Periodicity captures invariance amid variation. Periodicity generalization represents a model's ability to extract periodic patterns from training data and generalize to OOD scenarios. We introduce a unified interpretation of periodicity from the perspective of abstract algebra and reasoning, including both single and composite periodicity, to explain why Transformers struggle to generalize periodicity. Then we construct Coper about composite periodicity, a controllable generative benchmark with two OOD settings, Hollow and Extrapolation. Experiments reveal that periodicity generalization in Transformers is limited, where models can memorize periodic data during training, but cannot generalize to unseen composite periodicity. We release the source code to support future research.

cs.LG