SearcharxivSearch

arXiv subjects

Shijie Yuan

Publications and source records attributed to Shijie Yuan.

15 recordsLinked to original sources

A Multiplicative Fourier Proof of the Length-Four Index Conjecture

Let $C_n$ be a cyclic group of order $n$. We prove that if $(n,6)=1$, then every minimal zero-sum sequence of length four over $C_n$ has index one, thereby resolving the length-four index conjecture. After the gcd reduction, the nonunit case follows from the theorem of Shen-Xia-Li, and the remaining unit case is solved by a new multiplicative Fourier argument. The index-two residue identity yields a character-moment relation, and the odd characters with vanishing first moment form an exceptional spectrum of size at most $157φ(n)/1440<φ(n)/9$. A finite-group uncertainty principle then forces the four-term multiset to be invariant under negation, contradicting minimality. Apart from standard facts about primitive Dirichlet $L$-functions, the remaining argument is finite and requires neither asymptotic estimates nor computational verification.

math.NT

Robust Power and Sample Size Calculations in Quasi-likelihood Models: Methods and Practice

Accurate power and sample size (PSS) calculations are essential for designing studies that use quasi-likelihood (QL) models, which extend generalized linear models (GLMs) to settings where the full distribution of the outcome is not specified. Traditional PSS approaches often rely on restrictive distributional assumptions, limiting their applicability when responses have non-standard distributions, variance functions are misspecified, or when predictors exhibit complex dependence structures. Building on recent advances in effect size measures for PSS - specifically, 2 Standard Deviations in the Linear Predictor (2SLiP) and Pseudo-Partial $R^2$ (P2R2) - developed with interpretability in mind, this paper extends and evaluates these effect size measures in the QL framework, keying in particular on their utility in PSS. We assess their empirical performance for the Wald test and then extend to the score test through extensive simulations across diverse outcome types, link functions, and variance structures. To illustrate practical utility, we applied these effect size measures to survey data on frontline health care workers from \citet{cahill2022occupational} to quantify the association between perceived personal protective equipment adequacy and mental health outcomes during the COVID-19 pandemic, adjusting for covariates. Our findings demonstrate that both 2SLiP and P2R2 provide robust and interpretable alternatives to traditional methods, maintaining accuracy with minimal distributional assumptions and enhancing the flexibility of PSS for realistic study designs.

stat.ME

General measures of effect size to calculate power and sample size for Wald tests with generalized linear models

Power and sample size calculations for Wald tests in generalized linear models (GLMs) are often limited to specific cases like logistic regression. More general methods typically require detailed study parameters that are difficult to obtain during planning. We introduce two new effect size measures for estimating power and sample size in studies using Wald tests across any GLM. These measures accommodate any number of predictors or adjusters and require only basic study information. We provide practical guidance for interpreting and applying these measures to approximate a key parameter in power calculations. We also derive asymptotic bounds on the relative error of these approximations, showing that accuracy depends on features of the GLM such as the nonlinearity of the link function. To complement this analysis, we conduct simulation studies across common model specifications, identifying best use cases and opportunities for improvement. Finally, we test the methods in finite samples to confirm their practical utility, using a case study on the relationship between education and receipt of mental health treatment.

stat.ME

Monitoring Adverse Events Through Bayesian Nonparametric Clustering Across Studies

We introduce a Bayesian nonparametric inference approach for aggregate adverse event (AE) monitoring across studies. The proposed model seamlessly integrates external data from historical trials to define a relevant background rate and accommodates varying levels of covariate granularity (ranging from patient-level details to study-level aggregated summary data). Inference is based on a covariate-dependent product partition model (PPMx). A central element of the model is the ability to group experimental units with similar profiles. We introduce a pairwise similarity measure, with which we set up a random partition of experimental units with comparable covariate profiles, thereby improving the precision of AE rate estimation. Importantly, the proposed framework supports real-time safety monitoring under blinding with a seamless transition to unblinded analyses when indicated. Using one case study and simulation studies, we demonstrate the model's ability to detect safety signals and assess risk under diverse trial scenarios.

stat.ME

Reachable Sets-based Trajectory Planning Combining Reinforcement Learning and iLQR

The driving risk field is applicable to more complex driving scenarios, providing new approaches for safety decision-making and active vehicle control in intricate environments. However, existing research often overlooks the driving risk field and fails to consider the impact of risk distribution within drivable areas on trajectory planning, which poses challenges for enhancing safety. This paper proposes a trajectory planning method for intelligent vehicles based on the risk reachable set to further improve the safety of trajectory planning. First, we construct the reachable set incorporating the driving risk field to more accurately assess and avoid potential risks in drivable areas. Then, the initial trajectory is generated based on safe reinforcement learning and projected onto the reachable set. Finally, we introduce a trajectory planning method based on a constrained iterative quadratic regulator to optimize the initial solution, ensuring that the planned trajectory achieves optimal comfort, safety, and efficiency. We conduct simulation tests of trajectory planning in high-speed lane-changing scenarios. The results indicate that the proposed method can guarantee trajectory comfort and driving efficiency, with the generated trajectory situated outside high-risk boundaries, thereby ensuring vehicle safety during operation.

eess.SY

SIMBA -- A Bayesian Decision Framework for the Identification of Optimal Biomarker Subgroups for Cancer Basket Clinical Trials

We consider basket trials in which a biomarker-targeting drug may be efficacious for patients across different disease indications. Patients are enrolled if their cells exhibit some levels of biomarker expression. The threshold level is allowed to vary by indication. The proposed SIMBA method uses a decision framework to identify optimal biomarker subgroups (OBS) defined by an optimal biomarker threshold for each indication. The optimality is achieved through minimizing a posterior expected loss that balances estimation accuracy and investigator preference for broadly effective therapeutics. A Bayesian hierarchical model is proposed to adaptively borrow information across indications and enhance the accuracy in the estimation of the OBS. The operating characteristics of SIMBA are assessed via simulations and compared against a simplified version and an existing alternative method, both of which do not borrow information. SIMBA is expected to improve the identification of patient sub-populations that may benefit from a biomarker-driven therapeutics.

stat.AP

Human-aligned Safe Reinforcement Learning for Highway On-Ramp Merging in Dense Traffic

Most reinforcement learning (RL) approaches for the decision-making of autonomous driving consider safety as a reward instead of a cost, which makes it hard to balance the tradeoff between safety and other objectives. Human risk preference has also rarely been incorporated, and the trained policy might be either conservative or aggressive for users. To this end, this study proposes a human-aligned safe RL approach for autonomous merging, in which the high-level decision problem is formulated as a constrained Markov decision process (CMDP) that incorporates users' risk preference into the safety constraints, followed by a model predictive control (MPC)-based low-level control. The safety level of RL policy can be adjusted by computing cost limits of CMDP's constraints based on risk preferences and traffic density using a fuzzy control method. To filter out unsafe or invalid actions, we design an action shielding mechanism that pre-executes RL actions using an MPC method and performs collision checks with surrounding agents. We also provide theoretical proof to validate the effectiveness of the shielding mechanism in enhancing RL's safety and sample efficiency. Simulation experiments in multiple levels of traffic densities show that our method can significantly reduce safety violations without sacrificing traffic efficiency. Furthermore, due to the use of risk preference-aware constraints in CMDP and action shielding, we can not only adjust the safety level of the final policy but also reduce safety violations during the training stage, proving a promising solution for online learning in real-world environments.

cs.RO

The Modified Combo i3+3 Design for Novel-Novel Combination Dose-Finding Trials in Oncology

We consider a modified Ci3+3 (MCi3+3) design for dual-agent dose-finding trials in which both agents are tested on multiple doses. This usually happens when the agents are novel therapies. The MCi3+3 design offers a two-stage or three-stage version, depending on the practical need. The first stage begins with single-agent dose escalation, the second stage launches a model-free combination dose finding for both agents, and optionally, the third stage follows with a model-based design. MCi3+3 aims to maintain a relatively simple framework to facilitate practical application, while also address challenges that are unique to novel-novel combination dose finding. Through simulations, we demonstrate that the MCi3+3 design adeptly manages various toxicity scenarios. It exhibits operational characteristics on par with other combination designs, while offering an enhanced safety profile. The design is motivated and tested for a real-life clinical trial.

stat.AP

Pharmacometrics-Enabled DOse OPtimization (PEDOOP) for Seamless Phase I-II Trials in Oncology

We consider a dose-optimization design for first-in-human oncology trial that aims to identify a suitable dose for late-phase drug development. The proposed approach, called the Pharmacometrics-Enabled DOse OPtimization (PEDOOP) design, incorporates observed patient-level pharmacokinetics (PK) measurements and latent pharmacodynamics (PD) information for trial decision making and dose optimization. PEDOOP consists of two seamless phases. In phase I, patient-level time-course drug concentrations, derived PD effects, and the toxicity outcomes from patients are integrated into a statistical model to estimate the dose-toxicity response. A simple dose-finding design guides dose escalation in phase I. At the end of the phase I dose finding, a graduation rule is used to assess the safety and efficacy of all the doses and select those with promising efficacy and acceptable safety for a randomized comparison against a control arm in phase II. In phase II, patients are randomized to the selected doses based on a fixed or adaptive randomization ratio. At the end of phase II, an optimal biological dose (OBD) is selected for late-phase development. We conduct simulation studies to assess the PEDOOP design in comparison to an existing seamless design that also combines phases I and II in a single trial.

stat.AP

The Backfill i3+3 Design for Dose-Finding Trials in Oncology

We consider a formal statistical design that allows simultaneous enrollment of a main cohort and a backfill cohort of patients in a dose-finding trial. The goal is to accumulate more information at various doses to facilitate dose optimization. The proposed design, called Bi3+3, combines the simple dose-escalation algorithm in the i3+3 design and a model-based inference under the framework of probability of decisions (POD), both previously published. As a result, Bi3+3 provides a simple algorithm for backfilling patients to lower doses in a dose-finding trial once these doses exhibit safety profile in patients. The POD framework allows dosing decisions to be made when some backfill patients are still being followed with incomplete toxicity outcomes, thereby potentially expediting the clinical trial. At the end of the trial, Bi3+3 uses both toxicity and efficacy outcomes to estimate an optimal biological dose (OBD). The proposed inference is based on a dose-response model that takes into account either a monotone or plateau dose-efficacy relationship, which are frequently encountered in modern oncology drug development. Simulation studies show promising operating characteristics of the Bi3+3 design in comparison to existing designs.

stat.AP

A Unified Decision Framework for Phase I Dose-Finding Designs

The purpose of a phase I dose-finding clinical trial is to investigate the toxicity profiles of various doses for a new drug and identify the maximum tolerated dose. Over the past three decades, various dose-finding designs have been proposed and discussed, including conventional model-based designs, new model-based designs using toxicity probability intervals, and rule-based designs. We present a simple decision framework that can generate several popular designs as special cases. We show that these designs share common elements under the framework, such as the same likelihood function, the use of loss functions, and the nature of the optimal decisions as Bayes rules. They differ mostly in the choice of the prior distributions. We present theoretical results on the decision framework and its link to specific and popular designs like mTPI, BOIN, and CRM. These results provide useful insights into the designs and their underlying assumptions, and convey information to help practitioners select an appropriate design.

stat.ME

The Ci3+3 Design for Dual-Agent Combination Dose-Finding Clinical Trials

We propose a rule-based statistical design for combination dose-finding trials with two agents. The Ci3+3 design is an extension of the i3+3 design with simple decision rules comparing the observed toxicity rates and equivalence intervals that define the maximum tolerated dose combination. Ci3+3 consists of two stages to allow fast and efficient exploration of the dose-combination space. Statistical inference is restricted to a beta-binomial model for dose evaluation, and the entire design is built upon a set of fixed rules. We show via simulation studies that the Ci3+3 design exhibits similar and comparable operating characteristics to more complex designs utilizing model-based inferences. We believe that the Ci3+3 design may provide an alternative choice to help simplify the design and conduct of combination dose-finding trials in practice.

stat.AP

BaySize: Bayesian Sample Size Planning for Phase I Dose-Finding Trials

We propose BaySize, a sample size calculator for phase I clinical trials using Bayesian models. BaySize applies the concept of effect size in dose finding, assuming the MTD is defined based on an equivalence interval. Leveraging a decision framework that involves composite hypotheses, BaySize utilizes two prior distributions, the fitting prior (for model fitting) and sampling prior (for data generation), to conduct sample size calculation under desirable statistical power. Look-up tables are generated to facilitate practical applications. To our knowledge, BaySize is the first sample size tool that can be applied to a broad range of phase I trial designs.

stat.ME

Lessons Learned from the Bayesian Design and Analysis for the BNT162b2 COVID-19 Vaccine Phase 3 Trial

The phase III BNT162b2 mRNA COVID-19 vaccine trial is based on a Bayesian design and analysis, and the main evidence of vaccine efficacy is presented in Bayesian statistics. Confusion and mistakes are produced in the presentation of the Bayesian results. Some key statistics, such as Bayesian credible intervals, are mislabeled and stated as confidence intervals. Posterior probabilities of the vaccine efficacy are not reported as the main results. We illustrate the main differences in the reporting of Bayesian analysis results for a clinical trial and provide four recommendations. We argue that statistical evidence from a Bayesian trial, when presented properly, is easier to interpret and directly addresses the main clinical questions, thereby better supporting regulatory decision making. We also recommend using abbreviation "BI" to represent Bayesian credible intervals as a differentiation to "CI" which stands for confidence interval.

stat.AP

MUCE: Bayesian Hierarchical Modeling for the Design and Analysis of Phase 1b Multiple Expansion Cohort Trials

We propose a multiple cohort expansion (MUCE) approach as a design or analysis method for phase 1b multiple expansion cohort trials, which are novel first-in-human studies conducted following phase 1a dose escalation. The MUCE design is based on a class of Bayesian hierarchical models that adaptively borrow information across arms. Statistical inference is directly based on the posterior probability of each arm being efficacious, facilitating the decision making that decides which arm to select for further testing.

stat.ME