SearcharxivSearch

arXiv subjects

Ting Ye

Publications and source records attributed to Ting Ye.

At least 19 recordsLinked to original sources

An SPH--mesh Coupling for Vesicle Dynamics in Shear Flow

We present a novel computational framework that couples smoothed particle hydrodynamics~(SPH) with a triangulated membrane mesh to simulate the dynamics of vesicles suspended in fluids. A novel interface-tracking approach enforces membrane impermeability naturally, without resorting to non-physical constraints such as particle reflection or bounce-back boundary conditions. The membrane model incorporates four distinct bending energy formulations, namely the minimal model, the spontaneous curvature (SC) model, the bilayer couple (BC) model, and the area difference elasticity (ADE) model, providing a versatile tool for diverse biophysical scenarios. The framework is rigorously validated against equilibrium shapes and tank-treading motion of a vesicle, demonstrating excellent agreement with previous theoretical and numerical studies. A systematic investigation into the effects of each bending model on the vesicle's inclination angle, revolution frequency, and morphology in shear flow reveals key physical insights. Notably, spontaneous curvature has a negligible effect on steady-state orientation but profoundly alters rotational dynamics at low reduced volumes through the emergence of dumbbell-like shapes with deep constrictions. In contrast, the BC and ADE models induce characteristic asymmetric and stomatocyte morphologies. Our results establish the proposed SPH--mesh coupling as an accurate and robust tool for exploring the complex, shape-dependent dynamics of vesicles in fluid flows.

physics.flu-dyn

Shape-Preserving Covariate Adjustment via Empirical Likelihood in Randomized Experiment

Covariate adjustment improves estimation efficiency in randomized experiments, but standard calibration and augmentation methods, when applied to distribution or survival functions, do not preserve monotonicity---a fundamental property of the estimand. We propose using empirical likelihood with covariate-balancing constraints to construct a covariate-adjusted empirical measure for each treatment arm. Estimators of a broad class of distributional functionals, including cumulative distribution functions, survival functions, quantiles, and restricted mean survival times, are then derived as plug-in functionals of this measure, automatically inheriting proper shape constraints. We establish asymptotic normality with an explicit, guaranteed efficiency gain over unadjusted estimators. The asymptotic distributions are invariant to the randomization scheme, providing a unified inference procedure under simple randomization and all commonly used covariate-adaptive designs satisfying a mild balancing condition. This unified construction, adjusting the empirical measure once and deriving all estimators from it, offers a principled reconciliation of covariate adjustment with shape preservation. Simulations and an application to the SURPASS-4 trial confirm the theoretical gains.

stat.ME

Improving the efficiency of infectious disease prevention trials using negative control outcome event times

Baseline covariate adjustment can enhance the efficiency of randomized trials by improving precision of treatment effect estimates. However, the precision gain depends on how strongly the baseline covariates are prognostic for the primary outcome. In randomized trials of infectious disease prevention interventions (e.g., vaccines or passively administered antibodies), an individual's exposure to the pathogen is a leading prognostic factor but is rarely measurable at baseline. Hence, conventional covariate adjustment offers limited precision gain in prevention trials. We propose adjusting for a negative control outcome (NCO) event time, which is causally unaffected by the intervention but shares overlapping exposure mechanisms with the primary outcome. We formalize assumptions under which adjustment for the NCO event time is valid, and show that right-censoring of the NCO event time further complicates adjustment. We derive the efficient influence function for the treatment-arm-specific survivor function of the primary outcome when both the primary outcome and the NCO event time are right-censored, and use it to construct a cross-fitted, one-step estimator that is multiply robust to nuisance misspecification and asymptotically efficient when the nuisances are estimated accurately. In numerical experiments, our estimator compares comparably to benchmarks when the NCO event time is uninformative, and gains precision as the NCO event time is more prognostic for the primary outcome. We apply our method to HVTN 704/HPTN 085, a randomized, double-blinded trial of VRC01, a broadly neutralizing antibody against HIV-1. Adjusting for the time to a bacterial sexually transmitted infection --- a negative control outcome for HIV-1 acquisition --- reduced the estimated variance of the prevention efficacy estimate by approximately 27%, compared to roughly 2.5% for baseline covariate adjustment.

stat.ME

FlashTrie: A GPU-Accelerated Constrained Beam Search for Generative Retrieval

Constrained decoding is essential in generative retrieval, where document identifiers generated directly from a query must exactly match a predefined library of valid IDs. At scale, decoding is often constrained using a trie with beam search but most implementations run on CPU. Limited parallelism then makes trie traversal and candidate validation a serving bottleneck as beam width grows. We present FlashTrie, which addresses this limitation by optimizing constrained beam search on GPUs. It introduces an integer-aware succinct trie layout that uses bit compression to reduce memory footprint while keeping the full index in GPU high-bandwidth memory reducing memory stalls, and a cooperative CUDA kernel that performs beam expansion, validation, and pruning entirely on-device without per-step host orchestration. It further replaces CPU-style irregular lookup and heap maintenance with GPU-aware parallel primitives, improving warp utilization and reducing divergence. Together, these designs significantly reduce decoding latency and increase throughput while preserving retrieval quality. On a library of 800M keywords with beam widths up to 1000, FlashTrie reduces trie-search latency to under 3 ms, achieving up to 24x speedup over a highly optimized multi-threaded CPU baseline. These improvements enable FlashTrie to scale beam sizes by up to 5x in latency-critical applications such as sponsored search. In a large-scale online A/B experiment on a popular commercial search engine, it delivers a statistically significant +0.71% revenue lift, enabling real-time constrained decoding at a scale previously feasible only offline. The FlashTrie code will be publicly released after the review process.

cs.LG

Robust and Data-Adaptive Integration of Nonconcurrent Data in Platform Trials via Gaussian Processes

A platform trial is an innovative clinical trial design that enables simultaneous and continuous evaluation of multiple treatments within a single master protocol. Existing robust methods restrict analyses to concurrently randomized participants due to concerns that including nonconcurrent data may introduce bias from temporal trends. However, this exclusion represents a missed opportunity to improve efficiency. We propose a Gaussian process framework for incorporating nonconcurrent data that exploits temporal smoothness, a key feature of platform trials. The framework includes single-task and multi-task formulations and provides data-adaptive integration of nonconcurrent data with uncertainty quantification. The connection to kernel ridge regression yields a transparent frequentist interpretation of how nonconcurrent data are integrated. We establish two theoretical guarantees: incorporating nonconcurrent controls reduces the posterior variance of the treatment effect, and the resulting bias is controlled by a non-increasing bound. We extend the framework to discrete outcomes and to covariate adjustment, illustrate it on a hypothetical platform trial constructed from SURMOUNT-1, and provide an implementation in the R package RobinCID.

stat.ME

Evolving Longitudinal Patient Histories and Re-enrollment in Master Protocol Trials

A master protocol trial uses a single overarching protocol to test multiple therapies, often across several diseases or subtypes. Although such trials offer considerable flexibility and efficiency, their constrained and non-uniform treatment assignment raises two core challenges: precisely defining treatment effects and conducting robust, efficient inference. These challenges intensify when participants can re-enroll to receive additional eligible therapies over time. To address these issues, we first define a clinically meaningful estimand with a clear population specification for master protocol trials that allow re-enrollment across multiple episodes. Specifically, we define the episode-specific entire concurrently eligible (ECE) population, which preserves the integrity of randomized comparisons and remains invariant to randomization ratios and operational formats. We then introduce a per-episode added-effect estimand that aggregates episode-specific effects into an interpretable overall measure. For inference, we develop weighting and post-stratification estimators under the same minimal assumptions as conventional randomized trials, with model-assisted covariate adjustment to improve efficiency. We establish asymptotic distributions for all estimators and provide cluster-robust variance estimators that properly account for within-participant correlation induced by re-enrollment. We evaluate our methods through extensive simulations and apply our methods to SIMPLIFY, a master protocol trial comparing continuation versus discontinuation of two common cystic fibrosis therapies. All analyses are conducted using the \textsf{R} package \textsf{RobinCID}.

stat.ME

Proximal Learning for Trials With External Controls: A Case Study in HIV Prevention

With the advent of effective pre-exposure prophylaxis agents, active-controlled HIV prevention trials have become a common study design. Nevertheless, estimating absolute efficacy relative to a placebo remains important. In this paper, we introduce a novel application of proximal causal inference methods to estimate the counterfactual cumulative HIV incidence under placebo for participants in an active-controlled trial of cabotegravir, using external control data from a placebo-controlled trial with similar eligibility criteria. We leverage baseline sexually transmitted infection status and geographic region as negative control outcome and exposure variables, respectively. We address two key challenges: unmeasured differences in HIV risk between trials and statistical difficulties arising from low HIV incidence rates in both studies. To overcome these challenges, we develop two proximal inference approaches: (1) a semiparametric inverse probability of censoring weighting estimator, and (2) a two-stage regression-based strategy tailored to low-event-rate settings. Our theoretical and numerical investigations demonstrate these methods yield reliable estimates of the counterfactual one-year cumulative HIV incidence under placebo, and provide robust evidence of the superior efficacy of cabotegravir compared with placebo. These findings highlight the potential of proximal inference methods to estimate placebo-controlled effects in both single-arm and active-controlled trials by leveraging external controls.

stat.ME

Generative modeling for the bootstrap

Generative modeling builds on and substantially advances the classical idea of simulating synthetic data from observed samples. This paper shows that this principle is not only natural but also theoretically well-founded for bootstrap inference: it yields statistically valid confidence intervals that apply simultaneously to both regular and irregular estimators, including settings in which Efron's bootstrap fails. In this sense, the generative modeling-based bootstrap can be viewed as a modern version of the smoothed bootstrap: it could mitigate the curse of dimensionality and remain effective in challenging regimes where estimators may lack root-$n$ consistency or a Gaussian limit.

stat.ME

Covariate Adjustment for Wilcoxon Two Sample Statistic and Test

We apply covariate adjustment to the Wincoxon two sample statistic and Wincoxon-Mann-Whitney test in comparing two treatments. The covariate adjustment through calibration not only improves efficiency in estimation/inference but also widens the application scope of the Wilcoxon two sample statistic and Wincoxon-Mann-Whitney test to situations where covariate-adaptive randomization is used. We motivate how to adjust covariates to reduce variance, establish the asymptotic distribution of adjusted Wincoxon two sample statistic, and provide explicitly the guaranteed efficiency gain. The asymptotic distribution of adjusted Wincoxon two sample statistic is invariant to all commonly used covariate-adaptive randomization schemes so that a unified formula can be used in inference regardless of which covariate-adaptive randomization is applied.

stat.ME

The RobinCar Family: R Tools for Robust Covariate Adjustment in Randomized Clinical Trials

Purpose: Covariate adjustment is a powerful statistical technique that can increase efficiency in clinical trials. Recent guidance from the U.S. FDA provided recommendations and best practices for using covariate adjustment. However, there has existed a gap between the extensive statistical literature on covariate adjustment and software that is easy to use and abides by these best practices. Methods: We have developed the RobinCar Family, which is comprised of RobinCar and RobinCar2. These two R packages enable covariate-adjusted analyses for continuous, discrete, and time-to-event outcomes that follow best practices. For continuous and discrete outcomes, the functions in the RobinCar Family facilitate traditional forms of covariate adjustment such as ANCOVA as well as more recent approaches like ANHECOVA, G-computation with generalized linear models and machine learning models, and adjustment for a super-covariate (as in PROCOVA(TM)). Functions for time-to-event outcomes implement the covariate-adjusted log-rank test, the stratified covariate-adjusted log-rank test, and the marginal covariate-adjusted hazard ratio. The RobinCar Family is supported by the ASA Biopharmaceutical Section Covariate Adjustment Scientific Working Group. Results: We provide an accessible overview of the covariate-adjusted statistical methods, and describe how they are implemented in RobinCar and RobinCar2. We highlight important usage notes for clinical trial practitioners. Conclusion: We apply RobinCar and RobinCar2 functions by analyzing data from the AIDS Clinical Trials Group Study 175, demonstrating that they are straightforward and user-friendly.

stat.ME

ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates significant potential in enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing RLVR methods are often constrained by issues such as coarse-grained rewards, reward noise, and inefficient exploration, which lead to unstable training and entropy collapse. To address this challenge, we propose the Intrinsic Confidence-Driven Group Relative Preference Optimization method (ICPO). The intuition behind it lies in the fact that the probabilities of an LLM generating different responses can inherently and directly reflect its self-assessment of the reasoning process. Inspired by the idea of preference modeling, ICPO calculates a preference advantage score for each response by comparing the relative generation probabilities of multiple responses under the same input prompt, and integrates this score with verifiable rewards to guide the exploration process. We have discovered that the preference advantage score not only alleviates the issues of coarse-grained rewards and reward noise but also effectively curbs overconfident errors, enhances the relative superiority of undervalued high-quality responses, and prevents the model from overfitting to specific strategies. Comprehensive experiments across four general-domain benchmarks and three mathematical benchmarks demonstrate that ICPO steadily boosts reasoning compared to GRPO.

cs.AI

Debiasing hazard-based, time-varying vaccine effects using vaccine-irrelevant infections: An observational extension of a pivotal Phase 3 COVID-19 vaccine efficacy trial

Understanding how vaccine effectiveness (VE) changes over time can provide evidence-based guidance for public health decision making. While commonly reported by practitioners, time-varying VE estimates obtained using Cox regression are vul- nerable to hidden biases. To address these limitations, we describe how to leverage vaccine-irrelevant infections to identify hazard-based, time-varying VE in the pres- ence of unmeasured confounding and selection bias. We articulate assumptions under which our approach identifies a causal effect of an intervention deferring vaccination and interaction with the community in which infections circulate. We develop sieve and efficient influence curve-based estimators and discuss imposing monotone shape constraints and estimating VE against multiple variants. As a case study, we examine the observational booster phase of the Coronavirus Vaccine Efficacy (COVE) trial of the Moderna mRNA-1273 COVID-19 vaccine which used symptom-triggered multi- plex PCR testing to identify acute respiratory illnesses (ARIs) caused by SARS-CoV-2 and 20 off-target pathogens previously identified as compelling negative controls for COVID-19. Accounting for vaccine-irrelevant ARIs supported that the mRNA-1273 booster was more effective and durable against Omicron COVID-19 than suggested by Cox regression. Our work offers an approach to mitigate bias in hazard-based, time- varying treatment effects in randomized and non-randomized studies using negative controls.

stat.ME

Group Identification and Variable Selection in Multivariable Mendelian Randomization with Highly-Correlated Exposures

Multivariable Mendelian Randomization (MVMR) estimates the direct causal effects of multiple risk factors on an outcome using genetic variants as instruments. The growing availability of summary-level genetic data has created opportunities to apply MVMR in high-dimensional settings with many strongly correlated candidate risk factors. However, existing methods face three major limitations: weak instrument bias, limited interpretability, and the absence of valid post-selection inference. Here we introduce MVMR-PACS, a method that identifies signal-groups -- sets of causal risk factors with high genetic correlation or indistinguishable causal effects -- and estimates the direct effect of each group. MVMR-PACS minimizes a debiased objective function that reduces weak instrument bias while yielding interpretable estimates with theoretical guarantees for variable selection. We adapt a data-thinning strategy to summary-data MVMR to enable valid post-selection inference. In simulations, MVMR-PACS outperforms existing approaches in both estimation accuracy and variable selection. When applied to 27 lipoprotein subfraction traits and coronary artery disease risk, MVMR-PACS identifies biologically meaningful and robust signal-groups with interpretable direct causal effects.

stat.ME

Quasi Instrumental Variable Methods for Stable Hidden Confounding and Binary Outcome

Instrumental variable (IV) methods are central to causal inference from observational data, particularly when a randomized experiment is not feasible. However, of the three conventional core IV identification conditions, only one, IV relevance, is empirically verifiable; often one or both of the other conditions, exclusion restriction and IV independence from unmeasured confounders, are unmet in real-world applications. These challenges are compounded when the outcome is binary, a setting for which robust IV methods remain underdeveloped. A fundamental contribution of this paper is the development of a general identification strategy justified under a structural equilibrium dynamic generative model of so-called stable confounding and a quasi instrumental variable (QIV), i.e. a variable that is only assumed to be predictive of the outcome. Such a model implies (a) stability of confounding on the multiplicative scale, and (b) stability of the additive average treatment effect among the treated (ATT), across levels of that QIV. The former is all that is necessary to ensure a valid test of the causal null hypothesis; together those two conditions establish nonparametric identification and estimation of the conditional and marginal ATT. To address the statistical challenges posed by the need for boundedness in binary outcomes, we introduce a generalized odds product re-parametrization of the observed data distribution, and we develop both a principled maximum likelihood estimator and a triply robust semiparametric locally efficient estimator, which we evaluate through simulations and an empirical application to the UK Biobank.

stat.ME

GMM with Many Weak Moment Conditions and Nuisance Parameters: General Theory and Applications to Causal Inference

Weak identification arises in many statistical problems when key variables exhibit weak correlations-for example, when instrumental variables correlate weakly with treatment, or when proxy variables correlate weakly with unmeasured confounders. Under weak identification, standard estimation methods such as the generalized method of moments (GMM) can produce substantial bias, both in finite samples and asymptotically. This challenge is compounded in modern applications that require estimating many nuisance parameters. This paper develops a framework for estimation and inference of a finite-dimensional target parameter in general moment models with the number of weak moment conditions and nuisance parameters growing with sample size. We analyze a general two-step debiasing estimator that accommodates flexible, possibly nonparametric first-step estimation of nuisance parameters, in which Neyman orthogonality plays a more critical role in obtaining debiased inference than in conventional settings with strong identification. Under a many-weak-moment asymptotic regime, we establish the estimator's consistency and asymptotic normality. We provide high-level conditions for the general setting and demonstrate their application to two important special cases: inference with weak instruments and inference with weak proxies.

math.ST

High-rate continuous-variable quantum key distribution over 100 km fiber with composable security

Quantum key distribution (QKD), providing a way to generate secret keys with information-theoretic security,is arguably one of the most significant achievements in quantum information. The continuous-variable QKD (CV-QKD) offers the potential advantage of achieving a higher secret key rate (SKR) within a metro area, as well as being compatible with the mature telecom industry. However, the SKR and transmission distance of state-of-the-art CV-QKD systems are currently limited. Here, based on the novelly proposed orthogonal-frequency-division-multiplexing (OFDM) CV-QKD protocol, we demonstrate for the first time a high-rate multi-carrier (MC) CV-QKD with a 10 GHz symbol rate that chieves Gbps SKR within 10km and Mbps SKR over 100 km in the finite-size regime under composable security against collective attacks. The record-breaking results are achieved by suitable optimization of subcarrier number and modulation variance, well-controlled excess noise induced by both OFDM mechanism and efficient DSP scheme, and high-performance post-processing capacity realized by heterogeneous computing scheme. The composable finite-size SKR reaches 1779.45 Mbps@5km, 1025.49 Mbps@10km, 370.50 Mbps@25km, 99.93 Mbps@50km, 25.70 Mbps@75km,and 2.25 Mbps@100km, which improves the SKR by two orders of magnitude and quintuples the maximal transmission distance compared to most recently reported CV-QKD results [Nature Communications, 13, 4740 (2022)]. Interestingly, it is experimentally verified that the SKR of the proposed MC CV-QKD can approach five times larger than that of the single-carrier CV-QKD with the same symbol rate without additional hardware costs. Our work constitutes a critical step towards future high-speed quantum metropolitan and access networks.

quant-ph

Estimating treatment effects with competing intercurrent events in randomized controlled trials

The analysis of randomized controlled trials is often complicated by intercurrent events (IEs) -- events that occur after treatment initiation and affect either the interpretation or existence of outcome measurements. Examples include treatment discontinuation or the use of additional medications. In two recent clinical trials for systemic lupus erythematosus with complications of IEs, we classify the IEs into two broad categories: effect-informative (e.g., treatment discontinuation due to adverse events or lack of efficacy) and effect-uninformative (e.g., treatment discontinuation due to external factors such as pandemics or relocation). To define a clinically meaningful estimand, we adopt tailored strategies for each category of IEs. For effect-informative IEs, which are often informative about a patient's outcome, we use the composite variable strategy that assigns an outcome value indicative of treatment failure. For effect-uninformative IEs, we apply the hypothetical strategy, assuming their timing is conditionally independent of the outcome given treatment and baseline covariates, and hypothesizing a scenario in which such events do not occur. A central yet previously overlooked challenge is the presence of competing IEs, where the first IE censors all subsequent ones. Despite its ubiquity in practice, this issue has not been explicitly recognized or addressed in previous data analyses due to the lack of rigorous statistical methodology. In this paper, we propose a principled framework to formulate the estimand, establish its nonparametric identification and semiparametric estimation theory, and introduce weighting, outcome regression, and doubly robust estimators. We apply our methods to analyze the two systemic lupus erythematosus trials, demonstrating the robustness and practical utility of the proposed framework.

stat.ME

A sliced Wasserstein and diffusion approach to random coefficient models

We propose a new minimum-distance estimator for linear random coefficient models. This estimator integrates the recently advanced sliced Wasserstein distance with the nearest neighbor methods, both of which enhance computational efficiency. We demonstrate that the proposed method is consistent in approximating the true distribution. Moreover, our formulation naturally leads to a diffusion process-based algorithm and is closely connected to treatment effect distribution estimation -- both of which are of independent interest and hold promise for broader applications.

math.ST