SearcharxivSearch

arXiv subjects

Tianying Wang

Publications and source records attributed to Tianying Wang.

At least 19 recordsLinked to original sources

Augmented transfer regression learning for completely missing covariates

Large-scale population-level datasets, such as the UK Biobank and the All of Us Research Program, often lack covariates needed for a specific analysis, such as genetic or lifestyle measures, while related studies measure them. This creates a cross-population missing data problem in which covariates are completely unobserved in the target population, rather than partially missing within one dataset. We propose an augmented transfer regression learning method for this setting. The key identifying condition is a sub-population shift assumption: the joint distribution of the outcome and observed covariates may differ across source and target populations, but the conditional distribution of the missing covariates given observed variables is invariant. We combine importance-weighted estimating equations with imputation terms for first- and second-order moments of the missing covariates. The resulting estimator is doubly robust and remains consistent if either the density ratio model or both imputation models are correctly specified. It is n^{1/2}-consistent and asymptotically normal, and attains the semiparametric efficiency bound when both

stat.ME

Tree-aggregated compositional regression under measurement error

Compositional covariates in microbiome studies are often measured with error and organized by a biological hierarchy. Tree aggregation can improve multiresolution interpretation, but it also combines leaf-level errors into correlated contamination whose scale varies across the hierarchy. Existing tree aggregation and compositional measurement-error correction do not combine directly in redundant tree coordinates because a generic positive semidefinite projection can make the corrected criterion depend on the chosen representation. TARCO resolves this mismatch by normalizing tree coordinates by descendant leaf counts and applying a kernel-preserving positive semidefinite projection to the corrected tree-space Gram matrix. The resulting criterion is constant across equivalent tree representations, and the tree-coordinate estimator is exactly equivalent to an estimator on the identifiable coefficient space. For this estimator, we establish finite-sample prediction and coefficient-estimation bounds and, under sufficient separation and an appropriate grouping threshold, exact recovery of the maximal constant subtrees of the identifiable coefficient. These guarantees extend, with additional covariance-estimation terms, when the measurement-error covariance is estimated from independent auxiliary technical replicates. In a longitudinal gut microbiome analysis, correction changes the displayed taxonomic resolution of some associations with body mass index while preserving their directions, illustrating why the hierarchy should guide both signal aggregation and error correction.

stat.ME

Simulation-free extrapolation for misspecified models induced by categorizing an error-prone continuous covariate

Epidemiological studies often categorize continuous exposures for interpretation even when the underlying outcome-exposure association is continuous. The fitted categorical regression is a misspecified model because it replaces the continuous exposure with categories. With measurement error, categorization also misclassifies latent exposure categories, so the observed regression generally targets different means and contrasts. Existing estimating-equation and simulation-based extrapolation approaches require, respectively, outcome-model-specific derivations and pseudo-data generation with repeated fitting. We introduce simulation-free extrapolation (SIMFEX), which estimates misclassification probabilities and latent category proportions from replicates, computes the mean-scale trajectory without pseudo-data or repeated outcome-model fitting, and extrapolates it to the no-misclassification endpoint. Without additional error-free covariates, this construction applies across known one-to-one links, and the resulting estimator is consistent under stated conditions. With error-free covariates, the contrast relation remains exact for identity-link additive models when misclassification probabilities and category proportions are covariate-invariant; the nonidentity-link version provides a practical approximation. Simulations show substantial bias reduction relative to the naive analysis and coverage generally close to nominal. In the UK Biobank analysis, SIMFEX produced larger estimated high-versus-low fat-intake contrasts than the naive analysis for body mass index and obesity, illustrating how category misclassification can change the magnitude and uncertainty of prespecified contrasts.

stat.ME

Geometry of tail allocation in conformal prediction intervals

Lower and upper errors of a two-sided conformal prediction interval can have different scientific consequences. The division of target miscoverage between the two endpoints determines the corresponding tail-specific guarantees and can alter interval length at first order when tail scales differ. We characterize this allocation-length relation after separate one-sided split calibration, which preserves the tail-specific guarantees and marginal coverage whenever the allocation is selected independently of the calibration sample. Tail-quantile response to proportional rescaling determines the resulting length geometry. For regularly varying tails, normalized length converges to $g_γ(c)=c^{-ξ}+γ(1-c)^{-ξ}$, where $c$ is the upper-tail allocation fraction, $ξ$ is the tail index, and $γ$ is the lower-to-upper tail-scale ratio. A dominant tail produces a boundary optimum and makes the equal-tail interval asymptotically $2^ξ$ times as long as the optimum. Comparable tails produce an interior optimum, with equal-tail allocation optimal only at matching scales. An empirical allocation rule attains the corresponding optimum without estimating tail parameters. In the de Haan class the effect moves to an additive scale. Calibration resolution determines whether ordinary ranks can realize these allocations. When calibration tail counts remain bounded, two-sided rank feasibility also constrains the allocation. Tail homogeneity transfers the length relation over covariates, while opposite dominant tails preclude one globally efficient allocation.

stat.ME

Multi-Group Quadratic Discriminant Analysis via Projection

Multi-group classification arises in many prediction and decision-making problems, including applications in epidemiology, genomics, finance, and image recognition. Although classification methods have advanced considerably, much of the literature focuses on binary problems, and available extensions often provide limited flexibility for multi-group settings. Recent work has extended linear discriminant analysis to multiple groups, but more general methods are still needed to handle complex structures such as nonlinear decision boundaries and group-specific covariance patterns. We develop Multi-Group Quadratic Discriminant Analysis (MGQDA), a method for multi-group classification built on quadratic discriminant analysis. MGQDA projects high-dimensional predictors onto a lower-dimensional subspace, which enables accurate classification while capturing nonlinearity and heterogeneity in group-specific covariance structures. We derive theoretical guarantees, including variable selection consistency, to support the reliability of the procedure. In simulations and a gene-expression application, MGQDA achieves competitive or improved predictive performance compared with existing methods while selecting group-specific informative variables, indicating its practical value for high-dimensional multi-group classification problems. Supplementary materials for this article are available online.

stat.ME

A powerful transformation of quantitative responses for biobank-scale association studies

In linear regression models with non-Gaussian errors, transformations of the response variable are widely used in a broad range of applications. Motivated by various genetic association studies, transformation methods for hypothesis testing have received substantial interest. In recent years, the rise of biobank-scale genetic studies, which feature a vast number of participants that could be around half a million, spurred the need for new transformation methods that are both powerful for detecting weak genetic signals and computationally efficient for large-scale data. In this work, we propose a novel transformation method that leverages the information of the error density. This transformation leads to locally most powerful tests and therefore has strong power for detecting weak signals. To make the computation scalable to biobank-scale studies, we harnessed the nature of weak genetic signals and proposed a consistent and computationally efficient estimator of the transformation function. Through extensive simulations and a gene-based analysis of spirometry traits from the UK Biobank, we validate that our approach maintains stringent control over type I error rates and significantly enhances statistical power over existing methods.

stat.ME

A Semiparametric Quantile Single-Index Model for Zero-Inflated Outcomes

We consider the complex data modeling problem motivated by the zero-inflated and overdispersed data from microbiome studies. Analyzing how microbiome abundance is associated with human biological features, such as BMI, is of great importance for host health. Methods based on parametric distributional assumptions, such as zero-inflated Poisson and zero-inflated Negative Binomial regression, have been widely used in modeling such data, yet the parametric assumptions are restricted and hard to verify in real-world applications. We relax the parametric assumptions and propose a semiparametric single-index quantile regression model. It is flexible to include a wide range of possible association functions and adaptable to the various zero proportions across subjects, which relaxes the strong parametric distributional assumptions of most existing zero-inflated data modeling approaches. We establish the asymptotic properties for the index coefficients estimator and quantile regression curve estimation. Through extensive simulation studies, we demonstrate the superior performance of the proposed method regarding model fitting.

stat.ME

A unified quantile framework for nonlinear heterogeneous transcriptome-wide associations

Transcriptome-wide association studies (TWAS) are powerful tools for identifying gene-level associations by integrating genome-wide association studies and gene expression data. However, most TWAS methods focus on linear associations between genes and traits, ignoring the complex nonlinear relationships that may be present in biological systems. To address this limitation, we propose a novel framework, QTWAS, which integrates a quantile-based gene expression model into the TWAS model, allowing for the discovery of nonlinear and heterogeneous gene-trait associations. Via comprehensive simulations and applications to both continuous and binary traits, we demonstrate that the proposed model is more powerful than conventional TWAS in identifying gene-trait associations.

stat.ME

Debiased high-dimensional regression calibration for errors-in-variables log-contrast models

Motivated by the challenges in analyzing gut microbiome and metagenomic data, this work aims to tackle the issue of measurement errors in high-dimensional regression models that involve compositional covariates. This paper marks a pioneering effort in conducting statistical inference on high-dimensional compositional data affected by mismeasured or contaminated data. We introduce a calibration approach tailored for the linear log-contrast model. Under relatively lenient conditions regarding the sparsity level of the parameter, we have established the asymptotic normality of the estimator for inference. Numerical experiments and an application in microbiome study have demonstrated the efficacy of our high-dimensional calibration strategy in minimizing bias and achieving the expected coverage rates for confidence intervals. Moreover, the potential application of our proposed methodology extends well beyond compositional data, suggesting its adaptability for a wide range of research contexts.

stat.ME

ZIKQ: An innovative centile chart method for utilizing natural history data in rare disease clinical development

Utilizing natural history data as external control plays an important role in the clinical development of rare diseases, since placebo groups in double-blind randomization trials may not be available due to ethical reasons and low disease prevalence. This article proposed an innovative approach for utilizing natural history data to support rare disease clinical development by constructing reference centile charts. Due to the deterioration nature of certain rare diseases, the distributions of clinical endpoints can be age-dependent and have an absorbing state of zero, which can result in censored natural history data. Existing methods of reference centile charts can not be directly used in the censored natural history data. Therefore, we propose a new calibrated zero-inflated kernel quantile (ZIKQ) estimation to construct reference centile charts from censored natural history data. Using the application to Duchenne Muscular Dystrophy drug development, we demonstrate that the reference centile charts using the ZIKQ method can be implemented to evaluate treatment efficacy and facilitate a more targeted patient enrollment in rare disease clinical development.

stat.ME

Action and Trajectory Planning for Urban Autonomous Driving with Hierarchical Reinforcement Learning

Reinforcement Learning (RL) has made promising progress in planning and decision-making for Autonomous Vehicles (AVs) in simple driving scenarios. However, existing RL algorithms for AVs fail to learn critical driving skills in complex urban scenarios. First, urban driving scenarios require AVs to handle multiple driving tasks of which conventional RL algorithms are incapable. Second, the presence of other vehicles in urban scenarios results in a dynamically changing environment, which challenges RL algorithms to plan the action and trajectory of the AV. In this work, we propose an action and trajectory planner using Hierarchical Reinforcement Learning (atHRL) method, which models the agent behavior in a hierarchical model by using the perception of the lidar and birdeye view. The proposed atHRL method learns to make decisions about the agent's future trajectory and computes target waypoints under continuous settings based on a hierarchical DDPG algorithm. The waypoints planned by the atHRL model are then sent to a low-level controller to generate the steering and throttle commands required for the vehicle maneuver. We empirically verify the efficacy of atHRL through extensive experiments in complex urban driving scenarios that compose multiple tasks with the presence of other vehicles in the CARLA simulator. The experimental results suggest a significant performance improvement compared to the state-of-the-art RL methods.

cs.RO

Gaussian Processes with Errors in Variables: Theory and Computation

Covariate measurement error in nonparametric regression is a common problem in nutritional epidemiology and geostatistics, and other fields. Over the last two decades, this problem has received substantial attention in the frequentist literature. Bayesian approaches for handling measurement error have only been explored recently and are surprisingly successful, although the lack of a proper theoretical justification regarding the asymptotic performance of the estimators. By specifying a Gaussian process prior on the regression function and a Dirichlet process Gaussian mixture prior on the unknown distribution of the unobserved covariates, we show that the posterior distribution of the regression function and the unknown covariates density attain optimal rates of contraction adaptively over a range of Hölder classes, up to logarithmic terms. This improves upon the existing classical frequentist results which require knowledge of the smoothness of the underlying function to deliver optimal risk bounds. We also develop a novel surrogate prior for approximating the Gaussian process prior that leads to efficient computation and preserves the covariance structure, thereby facilitating easy prior elicitation. We demonstrate the empirical performance of our approach and compare it with competitors in a wide range of simulation experiments and a real data example.

math.ST

A Flexible Zero-Inflated Poisson-Gamma model with application to microbiome read counts

In microbiome studies, it is of interest to use a sample from a population of microbes, such as the gut microbiota community, to estimate the population proportion of these taxa. However, due to biases introduced in sampling and preprocessing steps, these observed taxa abundances may not reflect true taxa abundance patterns in the ecosystem. Repeated measures, including longitudinal study designs, may be potential solutions to mitigate the discrepancy between observed abundances and true underlying abundances. Yet, widely observed zero-inflation and over-dispersion issues can distort downstream statistical analyses aiming to associate taxa abundances with covariates of interest. To this end, we propose a Zero-Inflated Poisson Gamma (ZIPG) framework to address the aforementioned challenges. From a perspective of measurement errors, we accommodate the discrepancy between observations and truths by decomposing the mean parameter in Poisson regression into a true abundance level and a multiplicative measurement of sampling variability from the microbial ecosystem. Then, we provide a flexible model by connecting both mean abundance and the variability to different covariates, and build valid statistical inference procedures for both parameter estimation and hypothesis testing. Through comprehensive simulation studies and real data applications, the proposed ZIPG method provides significant insights into distinguished differential variability and abundance.

stat.ME

End-to-end Reinforcement Learning of Robotic Manipulation with Robust Keypoints Representation

We present an end-to-end Reinforcement Learning(RL) framework for robotic manipulation tasks, using a robust and efficient keypoints representation. The proposed method learns keypoints from camera images as the state representation, through a self-supervised autoencoder architecture. The keypoints encode the geometric information, as well as the relationship of the tool and target in a compact representation to ensure efficient and robust learning. After keypoints learning, the RL step then learns the robot motion from the extracted keypoints state representation. The keypoints and RL learning processes are entirely done in the simulated environment. We demonstrate the effectiveness of the proposed method on robotic manipulation tasks including grasping and pushing, in different scenarios. We also investigate the generalization capability of the trained model. In addition to the robust keypoints representation, we further apply domain randomization and adversarial training examples to achieve zero-shot sim-to-real transfer in real-world robotic manipulation tasks.

cs.RO

Deep N-ary Error Correcting Output Codes

Ensemble learning consistently improves the performance of multi-class classification through aggregating a series of base classifiers. To this end, data-independent ensemble methods like Error Correcting Output Codes (ECOC) attract increasing attention due to its easiness of implementation and parallelization. Specifically, traditional ECOCs and its general extension N-ary ECOC decompose the original multi-class classification problem into a series of independent simpler classification subproblems. Unfortunately, integrating ECOCs, especially N-ary ECOC with deep neural networks, termed as deep N-ary ECOC, is not straightforward and yet fully exploited in the literature, due to the high expense of training base learners. To facilitate the training of N-ary ECOC with deep learning base learners, we further propose three different variants of parameter sharing architectures for deep N-ary ECOC. To verify the generalization ability of deep N-ary ECOC, we conduct experiments by varying the backbone with different deep neural network architectures for both image and text classification tasks. Furthermore, extensive ablation studies on deep N-ary ECOC show its superior performance over other deep data-independent ensemble methods.

cs.CV

Integrated Quantile RAnk Test (iQRAT) for gene-level associations

Gene-based testing is a commonly employed strategy in many genetic association studies. Gene-trait associations can be complex due to underlying population heterogeneity, gene-environment interactions, and various other reasons. Existing gene-based tests, such as Burden and Sequence Kernel Association Tests (SKAT), are based on detecting differences in a single summary statistic, such as the mean or the variance, and may miss or underestimate higher-order associations that could be scientifically interesting. In this paper, we propose a new family of gene-level association tests which integrate quantile rank score processes to better accommodate complex associations. The resulting test statistics have multiple advantages: (1) they are almost as efficient as the best existing tests when the associations are homogeneous across quantile levels, and have improved efficiency for complex and heterogeneous associations, (2) they provide useful insights on risk stratification, (3) the test statistics are distribution-free, and could hence accommodate a wide range of underlying distributions, and (4) they are computationally efficient. We established the asymptotic properties of the proposed tests under the null and alternative hypothesis and conducted large scale simulation studies to investigate their finite sample performance. We applied the proposed tests to the Metabochip data to identify genetic associations with lipid traits and compared the results with those of the Burden and SKAT tests.

stat.ME

Improved Semiparametric Analysis of Polygenic Gene-Environment Interactions in Case-Control Studies

Standard logistic regression analysis of case-control data has low power to detect gene-environment interactions, but until recently it was the only method that could be used on complex polygenic data for which parametric distributional models are not feasible. Under the assumption of gene-environment independence in the underlying population, Stalder et al. (2017, Biometrika, 104, 801-812) developed a retrospective method that treats both genetic and environmental variables nonparametrically. However, the mathematical symmetry of genetic and environmental variables is overlooked. We propose an improvement to the method of Stalder et al. (2017) that increases the efficiency of the estimates with no additional assumptions and modest computational cost. This improvement is achieved by treating the genetic and environmental variables symmetrically to generate two sets of parameter estimates that are combined to generate a more efficient estimate. We employ a semiparametric framework to develop the asymptotic theory of the estimator, show its asymptotic efficiency gain, and evaluate its performance via simulation studies. The method is illustrated using data from a case-control study of breast cancer.

stat.ME

RoboCoDraw: Robotic Avatar Drawing with GAN-based Style Transfer and Time-efficient Path Optimization

Robotic drawing has become increasingly popular as an entertainment and interactive tool. In this paper we present RoboCoDraw, a real-time collaborative robot-based drawing system that draws stylized human face sketches interactively in front of human users, by using the Generative Adversarial Network (GAN)-based style transfer and a Random-Key Genetic Algorithm (RKGA)-based path optimization. The proposed RoboCoDraw system takes a real human face image as input, converts it to a stylized avatar, then draws it with a robotic arm. A core component in this system is the Avatar-GAN proposed by us, which generates a cartoon avatar face image from a real human face. AvatarGAN is trained with unpaired face and avatar images only and can generate avatar images of much better likeness with human face images in comparison with the vanilla CycleGAN. After the avatar image is generated, it is fed to a line extraction algorithm and converted to sketches. An RKGA-based path optimization algorithm is applied to find a time-efficient robotic drawing path to be executed by the robotic arm. We demonstrate the capability of RoboCoDraw on various face images using a lightweight, safe collaborative robot UR5.

cs.RO