SearcharxivSearch

arXiv subjects

Yanyuan Ma

Publications and source records attributed to Yanyuan Ma.

At least 19 recordsLinked to original sources

From matrix inversion to constraints: provably tighter confidence regions for importance weights in label shift

Importance weights are essential in domain adaptation under label shift, yet their utility is often undermined by the finite sample uncertainty associated with their estimation. Existing methods typically analyze this uncertainty through Gaussian elimination on interval-valued linear systems, which leads to overly conservative confidence regions and inefficient downstream applications. We propose a paradigm shift from inversion-based inference to a direct matrix constraint framework. We use this framework to define a joint confidence region and extract marginal intervals via linear programming, deriving provably tighter bounds for importance weights while maintaining exact finite-sample validity. Furthermore, we analyze the confidence region's geometry and provide the theoretical results for its diameter bounds. Evaluated across text, image, multimodal benchmarks, including AGNews, MNIST, CIFAR-10, N24News, and a real-world autonomous driving dataset, nuImages, our approach consistently yields shorter confidence intervals and smaller prediction sets than inversion-based methods.

stat.ML

PRESCCO: Efficient Prediction Intervals under a Right-Censored Covariate

In clinical studies, a patient's outcome, e.g., a cognitive test score, is typical or atypical depending on how it compares with the outcomes of patients at a similar point in a neurodegenerative disease, measured by how far they are from a common disease event. A prediction interval based on the time to that event provides that comparison, but that time is right-censored for most patients. Yet no method had computed such a prediction interval in this right-censored covariate setting. We first adapt three conformal prediction methods to this setting, and show that the estimated half-length, the distance the interval runs either side of its center, varies from one study to the next, so the same outcome can be judged typical in one study and atypical in another. We then develop the PRESCCO method, whose estimator of the half-length is semiparametrically efficient and doubly robust, staying consistent when one of its two models is misspecified, while the method loses no coverage. In simulations its standard deviation is four to thirty times smaller than under any conformal prediction method, and in a Huntington disease study with 77.2% right-censoring, an outcome typical in one study is no longer atypical in another.

stat.ME

Semiparametric Efficiency of Residual Correlation Testing under Gaussian Additive Noise Models

This paper studies conditional independence testing under the Gaussian additive noise model (GANM), where two variables are modeled as nonlinear functions of covariates with independent bivariate Gaussian regression errors. Under this framework, conditional independence can be characterized by the correlation coefficient of the regression errors, which motivates a test based on the Pearson correlation coefficient computed from the fitted residuals. Despite its simple form, the asymptotic behavior and statistical efficiency of the resulting test have not been well understood. In this paper, we develop the semiparametric efficiency theory under GANM and show, surprisingly, that the efficient estimator coincides exactly with the ordinary residual Pearson correlation estimator. We further establish the asymptotic properties of the proposed test and develop the corresponding inference procedure. Simulation studies demonstrate that the proposed method achieves near-oracle efficiency and competitive empirical power while maintaining valid Type I error control. We further apply the proposed test to conditional dependence analysis of U.S. stock returns.

math.ST

Calibrated Estimation and Inference for Semiparametric Regression Models

We consider a broad class of semiparametric regression models in which the conditional distribution of the response takes the form $f\{Y|\boldsymbol{x}^T\boldsymbolβ+m(z),ϕ\}$, known up to a parametric component $\boldsymbolβ$ of diverging dimension $p$, a smooth function $m(\cdot)$, and a dispersion parameter $ϕ$. The existing literature on such models has focused on semiparametric efficiency for $\boldsymbolβ$, treating $ϕ$ and $m(\cdot)$ as nuisances and largely ignoring finite-sample bias. Yet this bias can be substantial, particularly when $p$ is large relative to $n$ or the dispersion is high, and it can seriously undermine inference for $\boldsymbolβ$; moreover, $ϕ$ is often of direct scientific interest. We therefore propose SABRE, a general calibration framework for semiparametric estimation and inference, which calibrates an initial estimator against its model-implied expectation under a tractable parametric approximation to the semiparametric model. For generalized partially linear models, we show that SABRE reduces the bias of both $\boldsymbolβ$ and $ϕ$, accommodates a diverging parameter dimension without sparsity, and preserves the first-order variance and semiparametric efficiency of the initial estimator; the joint construction also improves estimation and inference for $m(\cdot)$. Simulation studies and an application to Alzheimer's disease genetics association analysis demonstrate the empirical effectiveness of SABRE in reducing bias and improving inference.

stat.ME

SPYCE: A Doubly Robust Estimator for Trials Targeting Early Huntington Disease under Outcome-Dependent Censoring

Clinical trials for neurodegenerative diseases must identify sensitive endpoints -- outcomes that change rapidly enough to detect treatment effects. In Huntington disease, this requires measuring how outcomes change as participants approach Stage 1. Yet many participants exit studies before reaching this stage, making their time to Stage 1 right-censored. Estimating how outcomes change requires models for both time to Stage 1 and time to study exit. When participants with worse outcomes exit earlier, this outcome-dependent censoring causes existing estimators to produce contradictory results: for the same cognitive outcome, one estimator suggests improvement while another shows decline. Existing estimators either ignore outcome-dependent censoring or require one model to be correctly specified, with no protection when it is not. We introduce SPYCE, a doubly robust estimator (consistent when either model is correctly specified) that achieves the smallest possible variance and allows both models to be estimated nonparametrically without sacrificing efficiency. Applied to data from PREDICT-HD, an observational Huntington disease study, SPYCE resolves current contradictions, identifies caudate and putamen volume ratios as the most promising sensitive endpoints, and shows that as few as 241 participants per arm are needed to detect treatment effects, versus hundreds of thousands under estimators that cannot handle outcome-dependent censoring.

stat.ME

Quantile regression with measurement errors

We devise a novel estimator for a general quantile regression model with normal measurement errors in the covariates. The method is applicable to both linear and nonlinear quantile regressions and does not impose the quantile requirement on multiple quantile levels simultaneously. We circumvent the difficulties caused by discontinuity in quantile regression through kernel smoothing, and overcome the nonlinearity inherent in quantile regression via considering extension to the complex domain and moment generating functions. We show that the resulting estimator achieves the standard root-$n$ consistency and asymptotic normality under mild conditions. The performance of the proposed method is illustrated via numerical simulations and a real data example related to Cherry Blossom times in Japan in 2024. This is the first consistent estimator in a general quantile regression problem with normal measurement errors.

stat.ME

Matrix-valued Network Autoregression Model with Latent Group Structure

Matrix-valued time series data are frequently observed in a broad range of areas and have attracted great attention recently. In this work, we model network effects for high dimensional matrix-valued time series data in a matrix autoregression framework. To characterize the potential heterogeneity of the subjects and handle the high dimensionality simultaneously, we assume that each subject has a latent group label, which enables us to cluster the subject into the corresponding row and column groups. We propose a group matrix network autoregression (GMNAR) model, which assumes that the subjects in the same group share the same set of model parameters. To estimate the model, we develop an iterative algorithm. Theoretically, we show that the group-wise parameters and group memberships can be consistently estimated when the group numbers are correctly or possibly over-specified. An information criterion for group number estimation is also provided to consistently select the group numbers. Lastly, we implement the method on a Yelp dataset to illustrate the usefulness of the method.

stat.ME

Multi-relational Network Autoregression Model with Latent Group Structures

Multi-relational networks among entities are frequently observed in the era of big data. Quantifying the effects of multiple networks have attracted significant research interest recently. In this work, we model multiple network effects through an autoregressive framework for tensor-valued time series. To characterize the potential heterogeneity of the networks and handle the high dimensionality of the time series data simultaneously, we assume a separate group structure for entities in each network and estimate all group memberships in a data-driven fashion. Specifically, we propose a group tensor network autoregression (GTNAR) model, which assumes that within each network, entities in the same group share the same set of model parameters, and the parameters differ across networks. An iterative algorithm is developed to estimate the model parameters and the latent group memberships simultaneously. Theoretically, we show that the group-wise parameters and group memberships can be consistently estimated when the group numbers are correctly- or possibly over-specified. An information criterion for group number estimation of each network is also provided to consistently select the group numbers. Lastly, we implement the method on a Yelp dataset to illustrate the usefulness of the method.

stat.ME

Complier General Causal Effect in Randomized Controlled Trials with One-Sided Noncompliance

A randomized controlled trial (RCT) is widely regarded as the gold standard for assessing the causal effect of a treatment or intervention, assuming perfect implementation. In practice, however, randomization can be compromised for various reasons, such as one-sided noncompliance. In this paper, we first systematically study the likelihood-based identifiability in an RCT with one-sided noncompliance. This foundational analysis naturally gives rise to the complier general causal effect (CGCE) as the primary estimand. We further develop two estimators for the CGCE: a simple estimator that requires no nonparametric procedures, and an efficient estimator that achieves the semiparametric efficiency bound. Our theoretical analysis shows that, achieving semiparametric efficiency requires only the nuisance estimators to converge in $L_2$-norm, with no restriction on their convergence rates. This rate-free property opens the door to employing many more modern machine learning methods while still guaranteeing efficiency. Comprehensive simulation studies and a real data application are conducted to illustrate the proposed methods and to compare them with existing approaches.

stat.ME

Unsupervised Domain Adaptation for Binary Classification with an Unobservable Source Subpopulation

We study an unsupervised domain adaptation problem where the source domain consists of subpopulations defined by the binary label $Y$ and a binary background (or environment) $A$. We focus on a challenging setting in which one such subpopulation in the source domain is unobservable. Naively ignoring this unobserved group can result in biased estimates and degraded predictive performance. Despite this structured missingness, we show that the prediction in the target domain can still be recovered. Specifically, we rigorously derive both background-specific and overall prediction models for the target domain. For practical implementation, we propose the distribution matching method to estimate the subpopulation proportions. We provide theoretical guarantees for the asymptotic behavior of our estimator, and establish an upper bound on the prediction error. Experiments on both synthetic and real-world datasets show that our method outperforms the naive benchmark that does not account for this unobservable source subpopulation.

stat.ML

Borrowing Information from an Unidentifiable Model: Guaranteed Efficiency Gain with a Dichotomized Outcome in the External Data

In the era of big data, the increasing availability of diverse data sources has driven interest in analytical approaches that integrate information across sources to enhance statistical accuracy, efficiency, and scientific insights. Many existing methods assume exchangeability among data sources and often implicitly require that sources measure identical covariates or outcomes, or that the error distribution is correctly specified-assumptions that may not hold in complex real-world scenarios. This paper explores the integration of data from sources with distinct outcome scales, focusing on leveraging external data to improve statistical efficiency. Specifically, we consider a scenario where the primary dataset includes a continuous outcome, and external data provides a dichotomized version of the same outcome. We propose two novel estimators: the first estimator remains asymptotically consistent even when the error distribution is potentially misspecified, while the second estimator guarantees an efficiency gain over weighted least squares estimation that uses the primary study data alone. Theoretical properties of these estimators are rigorously derived, and extensive simulation studies are conducted to highlight their robustness and efficiency gains across various scenarios. Finally, a real-world application using the NHANES dataset demonstrates the practical utility of the proposed methods.

stat.ME

A Communication-Efficient Distributed Algorithm for Learning with Heterogeneous and Structurally Incomplete Multi-Site Data

In multicenter biomedical research, integrating data from multiple decentralized sites provides more robust and generalizable findings due to its larger sample size and the ability to account for the between-site heterogeneity. However, sharing individual-level data across sites is often difficult due to patient privacy concerns and regulatory restrictions. To overcome this challenge, many distributed algorithms, that fit a global model by only communicating aggregated information across sites, have been proposed. A major challenge in applying existing distributed algorithms to real-world data is that their validity often relies on the assumption that data across sites are independently and identically distributed, which is frequently violated in practice. In biomedical applications, data distributions across clinical sites can be heterogeneous. Additionally, the set of covariates available at each site may vary due to different data collection protocols. We propose a distributed inference framework for data integration in the presence of both distribution heterogeneity and data structural heterogeneity. By modeling heterogeneous and structurally missing data using density-tilted generalized method of moments, we developed a general aggregated data-based distributed algorithm that is communication-efficient and heterogeneity-aware. We establish the asymptotic properties of our estimator and demonstrate the validity of our method via simulation studies.

stat.ME

Robust Estimation under Outcome Dependent Right Censoring in Huntington Disease: Estimators for Low and High Censoring Rates

Across health applications, researchers model outcomes as a function of time to an event, but the event time is right-censored for participants who exit the study or otherwise do not experience the event during follow-up. When censoring depends on the outcome-as in neurodegenerative disease studies where dropout is potentially related to disease severity-standard regression estimators produce biased estimates. We develop three consistent estimators for this outcome-dependent censoring setting: two augmented inverse probability weighted (AIPW) estimators and one maximum likelihood estimator (MLE). We establish their asymptotic properties and derive their robust sandwich variance estimators that account for nuisance parameter estimation. A key contribution is demonstrating that the choice of estimator to use depends on the censoring rate-the MLE performs best under low censoring rates, while the AIPW estimators yield lower bias and a higher nominal coverage under high censoring rates. We apply our estimators to Huntington disease data to characterize health decline leading up to mild cognitive impairment onset. The AIPW estimator with robustness matrix provided clinically-backed estimates with improved precision over inverse probability weighting, while MLE exhibited bias. Our results provide practical guidance for estimator selection based on censoring rate.

stat.ME

Super doubly robust and efficient estimator for informative covariate censoring

Early intervention in neurodegenerative diseases requires identifying periods before diagnosis when decline is rapid enough to detect whether a therapy is slowing progression. Since rapid decline typically occurs close to diagnosis, identifying these periods requires knowing each patient's time of diagnosis. Yet many patients exit studies before diagnosis, making time of diagnosis right-censored by time of study exit -- creating a right-censored covariate problem when estimating decline. Existing estimators either assume noninformative covariate censoring, where time of study exit is independent of time of diagnosis, or allow informative covariate censoring, but require correctly specifying how these times are related. We developed SPIRE (Semi-Parametric Informative Right-censored covariate Estimator), a super doubly robust estimator that remains consistent without correctly specifying densities governing time of diagnosis or time of study exit. Typical double robustness requires at least one density to be correct; SPIRE requires neither. When both densities are correctly specified, SPIRE achieves semiparametric efficiency. We also developed a test for detecting informative covariate censoring. Simulations with 85% right-censoring demonstrated SPIRE's robustness, efficiency and reliable detection of informative covariate censoring. Applied to Huntington disease data, SPIRE handled informative covariate censoring appropriately and remained consistent regardless of density specification, providing a reliable tool for early intervention.

math.ST

A KL-divergence based test for elliptical distribution

We conduct a KL-divergence based procedure for testing elliptical distributions. The procedure simultaneously takes into account the two defining properties of an elliptically distributed random vector: independence between length and direction, and uniform distribution of the direction. The test statistic is constructed based on the $k$ nearest neighbors ($k$NN) method, and two cases are considered where the mean vector and covariance matrix are known and unknown. First-order asymptotic properties of the test statistic are rigorously established by creatively utilizing sample splitting, truncation and transformation between Euclidean space and unit sphere, while avoiding assuming Fréchet differentiability of any functionals. Debiasing and variance inflation are further proposed to treat the degeneration of the influence function. Numerical implementations suggest better size and power performance than the state of the art procedures.

stat.ME

Efficient Inference under Label Shift in Unsupervised Domain Adaptation

In many real-world applications, researchers aim to deploy models trained in a source domain to a target domain, where obtaining labeled data is often expensive, time-consuming, or even infeasible. While most existing literature assumes that the labeled source data and the unlabeled target data follow the same distribution, distribution shifts are common in practice. This paper focuses on label shift and develops efficient inference procedures for general parameters characterizing the unlabeled target population. A central idea is to model the outcome density ratio between the labeled and unlabeled data. To this end, we propose a progressive estimation strategy that unfolds in three stages: an initial heuristic guess, a consistent estimation, and ultimately, an efficient estimation. This self-evolving process is novel in the statistical literature and of independent interest. We also highlight the connection between our approach and prediction-powered inference (PPI), which uses machine learning models to improve statistical inference in related settings. We rigorously establish the asymptotic properties of the proposed estimators and demonstrate their superior performance compared to existing methods. Through simulation studies and multiple real-world applications, we illustrate both the theoretical contributions and practical benefits of our approach.

stat.ME

Bioequivalence Assessment for Locally Acting Drugs: A Framework for Feasible and Efficient Evaluation

Equivalence testing plays a key role in several domains, such as the development of generic medical products, which are therapeutically equivalent to brand-name drugs but with reduced cost and increased accessibility. Promoting access to generics is a critical public health issue with substantial societal implications, but establishing equivalence is particularly challenging in multivariate settings. A notable example refers to locally acting drugs designed to exert their therapeutic effects at a localized area where they are administered rather than being absorbed into the bloodstream, where complex experimental protocols lead to reduced sample sizes and substantial experimental noise. Traditional approaches, such as the Two One-Sided Tests (TOST), cannot adequately tackle the complex multivariate nature of such data. In this work, we develop an adjustment for the TOST procedure by simultaneously correcting its significance level and equivalence margins to ensure control of the test size and increase its power. In large samples, this approach leads to an optimal adjustment for the univariate TOST procedure. In multivariate settings, where we show that an optimal adjustment does not exist, our proposal maintains equal marginal test sizes and overall size control while maximizing power in important cases. Through extensive simulation studies and a case study on multivariate bioequivalence assessment for two antifungal topical products, we demonstrate the superior performance of our method across various scenarios encountered in practice.

stat.ME

Survival analysis under label shift

Let P represent the source population with complete data, containing covariate $\mathbf{Z}$ and response $T$, and Q the target population, where only the covariate $\mathbf{Z}$ is available. We consider a setting with both label shift and label censoring. Label shift assumes that the marginal distribution of $T$ differs between $P$ and $Q$, while the conditional distribution of $\mathbf{Z}$ given $T$ remains the same. Label censoring refers to the case where the response $T$ in $P$ is subject to random censoring. Our goal is to leverage information from the label-shifted and label-censored source population $P$ to conduct statistical inference in the target population $Q$. We propose a parametric model for $T$ given $\mathbf{Z}$ in $Q$ and estimate the model parameters by maximizing an approximate likelihood. This allows for statistical inference in $Q$ and accommodates a range of classical survival models. Under the label shift assumption, the likelihood depends not only on the unknown parameters but also on the unknown distribution of $T$ in $P$ and $\mathbf{Z}$ in $Q$, which we estimate nonparametrically. The asymptotic properties of the estimator are rigorously established and the effectiveness of the method is demonstrated through simulations and a real data application. This work is the first to combine survival analysis with label shift, offering a new research direction in this emerging topic.

stat.ME