SearcharxivSearch

arXiv subjects

Haiming Zhou

Publications and source records attributed to Haiming Zhou.

12 recordsLinked to original sources

BOIN Designs for Dose Escalation With Selected Dose Combinations in Oncology Phase I Trials

In phase I dose escalation studies for dual-agent combinations, at least one drug often has an established monotherapy dose. Consequently, substantial prior clinical safety data often exist for one or more monotherapies, allowing the study to focus on a subset of selected dose combinations rather than exhaustively evaluating all possible dose combinations for two agents. The Bayesian Optimal Interval (BOIN) design framework is widely recognized for its robust performance and ease of implementation; however, the BOIN for combination design, abbreviated as BOIN-C in this paper, was originally developed to evaluate full combinations and may not be directly applicable for the subset of selected combinations. In this paper, we propose three extensions to the BOIN-C design to address scenarios involving selected dose combinations: (a) BOIN-CS: a generalized BOIN-C design to accommodate any subset of dose combinations. (b) BOIN-CE: Exploration of new off-diagonal dose combinations when de-escalating. This option provides additional opportunities to treat patients with dose combinations that have not been administered. (c) BOIN-CB: Bayesian logistic regression model (BLRM)-guided BOIN design, which uses the BLRM model to break the tie when two dose combinations have an equal posterior probability of being selected. This can be useful when the dose-toxicity relationship is expected to be reasonably aligned with a logistic relationship. These study design options are motivated by practical considerations, and their operating characteristics are evaluated through extensive simulations under various scenarios, demonstrating satisfactory performance.

stat.ME

A Bayesian Basket Trial Design Using Local Power Prior

In recent years, basket trials, which allow the evaluation of an experimental therapy across multiple tumor types within a single protocol, have gained prominence in early-phase oncology development. Unlike traditional trials, which evaluate each tumor type separately and often face challenges with limited sample sizes, basket trials offer the advantage of borrowing information across various tumor types to enhance statistical power. However, a key challenge in designing basket trials is determining the appropriate extent of information borrowing while maintaining an acceptable type I error rate control. In this paper, we propose a novel 3-component local power prior (local-PP) framework that introduces a dynamic and flexible approach to information borrowing. The framework consists of three components: global borrowing control, pairwise similarity assessments, and a borrowing threshold, allowing for tailored and interpretable borrowing across heterogeneous tumor types. Unlike many existing Bayesian methods that rely on computationally intensive Markov chain Monte Carlo (MCMC) sampling, the proposed approach provides a closed-form solution, significantly reducing computation time in large-scale simulations for evaluating operating characteristics. Extensive simulations demonstrate that the proposed local-PP framework performs comparably to more complex methods while significantly shortening computation time.

stat.AP

The flexible Gumbel distribution: A new model for inference about the mode

A new unimodal distribution family indexed by the mode and three other parameters is derived from a mixture of a Gumbel distribution for the maximum and a Gumbel distribution for the minimum. Properties of the proposed distribution are explored, including model identifiability and flexibility in capturing heavy-tailed data that exhibit different directions of skewness over a wide range. Both frequentist and Bayesian methods are developed to infer parameters in the new distribution. Simulation studies are conducted to demonstrate satisfactory performance of both methods. By fitting the proposed model to simulated data and data from an application in hydrology, it is shown that the proposed flexible distribution is especially suitable for data containing extreme values in either direction, with the mode being a location parameter of interest. Using the proposed unimodal distribution, one can easily formulate a regression model concerning the mode of a response given covariates. We apply this model to data from an application in criminology to reveal interesting data features that are obscured by outliers. Computer programs for implementing all considered inference methods in the study are available at https://github.com/rh8liuqy/flexible_Gumbel.

stat.ME

Early Indicators of Scientific Impact: Predicting Citations with Altmetrics

Identifying important scholarly literature at an early stage is vital to the academic research community and other stakeholders such as technology companies and government bodies. Due to the sheer amount of research published and the growth of ever-changing interdisciplinary areas, researchers need an efficient way to identify important scholarly work. The number of citations a given research publication has accrued has been used for this purpose, but these take time to occur and longer to accumulate. In this article, we use altmetrics to predict the short-term and long-term citations that a scholarly publication could receive. We build various classification and regression models and evaluate their performance, finding neural networks and ensemble models to perform best for these tasks. We also find that Mendeley readership is the most important factor in predicting the early citations, followed by other factors such as the academic status of the readers (e.g., student, postdoc, professor), followers on Twitter, online post length, author count, and the number of mentions on Twitter, Wikipedia, and across different countries.

cs.DL

Parametric mode regression for bounded responses

We propose new parametric frameworks of regression analysis with the conditional mode of a bounded response as the focal point of interest. Covariate effects estimation and prediction based on the maximum likelihood method under two new classes of regression models are demonstrated. We also develop graphical and numerical diagnostic tools to detect various sources of model misspecification. Predictions based on different central tendency measures inferred using various regression models are compared using synthetic data in simulations. Finally, we conduct regression analysis for data from the Alzheimer's Disease Neuroimaging Initiative to demonstrate practical implementation of the proposed methods. Supplementary materials that contain technical details, and additional simulation and data analysis results are available online.

stat.ME

Conditional density estimation with covariate measurement error

We consider estimating the density of a response conditioning on an error-prone covariate. Motivated by two existing kernel density estimators in the absence of covariate measurement error, we propose a method to correct the existing estimators for measurement error. Asymptotic properties of the resultant estimators under different types of measurement error distributions are derived. Moreover, we adjust bandwidths readily available from existing bandwidth selection methods developed for error-free data to obtain bandwidths for the new estimators. Extensive simulation studies are carried out to compare the proposed estimators with naive estimators that ignore measurement error, which also provide empirical evidence for the effectiveness of the proposed bandwidth selection methods. A real-life data example is used to illustrate implementation of these methods under practical scenarios. An R package, lpme, is developed for implementing all considered methods, which we demonstrate via an R code example in Appendix H.

stat.ME

spBayesSurv: Fitting Bayesian Spatial Survival Models Using R

Spatial survival analysis has received a great deal of attention over the last 20 years due to the important role that geographical information can play in predicting survival. This paper provides an introduction to a set of programs for implementing some Bayesian spatial survival models in R using the package spBayesSurv. The function survregbayes includes the three most commonly-used semiparametric models: proportional hazards, proportional odds, and accelerated failure time. All manner of censored survival times are simultaneously accommodated including uncensored, interval censored, current-status, left and right censored, and mixtures of these. Left-truncated data are also accommodated. Time-dependent covariates are allowed under the piecewise constant assumption. Both georeferenced and areally observed spatial locations are handled via frailties. Model fit is assessed with conditional Cox-Snell residual plots, and model choice is carried out via the log pseudo marginal likelihood, the deviance information criterion and the Watanabe-Akaike information criterion. The accelerated failure time frailty model with a covariate-dependent baseline is included in the function frailtyGAFT. In addition, the package also provides two marginal survival models: proportional hazards and linear dependent Dirichlet process mixture, where the spatial dependence is modeled via spatial copulas. Note that the package can also handle non-spatial data using non-spatial versions of aforementioned models.

stat.CO

Bandwidth selection for nonparametric modal regression

In the context of estimating local modes of a conditional density based on kernel density estimators, we show that existing bandwidth selection methods developed for kernel density estimation are unsuitable for mode estimation. We propose two methods to select bandwidths tailored for mode estimation in the regression setting. Numerical studies using synthetic data and a real-life data set are carried out to demonstrate the performance of the proposed methods in comparison with several well received bandwidth selection methods for density estimation.

stat.CO

A unified framework for fitting Bayesian semiparametric models to arbitrarily censored survival data, including spatially-referenced data

A comprehensive, unified approach to modeling arbitrarily censored spatial survival data is presented for the three most commonly-used semiparametric models: proportional hazards, proportional odds, and accelerated failure time. Unlike many other approaches, all manner of censored survival times are simultaneously accommodated including uncensored, interval censored, current-status, left and right censored, and mixtures of these. Left-truncated data are also accommodated leading to models for time-dependent covariates. Both georeferenced (location exactly observed) and areally observed (location known up to a geographic unit such as a county) spatial locations are handled; formal variable selection makes model selection especially easy. Model fit is assessed with conditional Cox-Snell residual plots, and model choice is carried out via LPML and DIC. Baseline survival is modeled with a novel transformed Bernstein polynomial prior. All models are fit via a new function which calls efficient compiled C++ in the R package spBayesSurv. The methodology is broadly illustrated with simulations and real data applications. An important finding is that proportional odds and accelerated failure time models often fit significantly better than the commonly-used proportional hazards model. Supplementary materials are available online.

stat.AP

An alternative local polynomial estimator for the error-in-variables problem

We consider the problem of estimating a regression function when a covariate is measured with error. Using the local polynomial estimator of Delaigle, Fan, and Carroll (2009) as a benchmark, we propose an alternative way of solving the problem without transforming the kernel function. The asymptotic properties of the alternative estimator are rigorously studied. A detailed implementing algorithm and a computationally efficient bandwidth selection procedure are also provided. The proposed estimator is compared with the existing local polynomial estimator via extensive simulations and an application to the motorcycle crash data. The results show that the new estimator can be less biased than the existing estimator and is numerically more stable.

stat.ME

Nonparametric modal regression in the presence of measurement error

In the context of regressing a response $Y$ on a predictor $X$, we consider estimating the local modes of the distribution of $Y$ given $X=x$ when $X$ is prone to measurement error. We propose two nonparametric estimation methods, with one based on estimating the joint density of $(X, Y)$ in the presence of measurement error, and the other built upon estimating the conditional density of $Y$ given $X=x$ using error-prone data. We study the asymptotic properties of each proposed mode estimator, and provide implementation details including the mean-shift algorithm for mode seeking and bandwidth selection. Numerical studies are presented to compare the proposed methods with an existing mode estimation method developed for error-free data naively applied to error-prone data.

stat.ME

Modeling county level breast cancer survival data using a covariate-adjusted frailty proportional hazards model

Understanding the factors that explain differences in survival times is an important issue for establishing policies to improve national health systems. Motivated by breast cancer data arising from the Surveillance Epidemiology and End Results program, we propose a covariate-adjusted proportional hazards frailty model for the analysis of clustered right-censored data. Rather than incorporating exchangeable frailties in the linear predictor of commonly-used survival models, we allow the frailty distribution to flexibly change with both continuous and categorical cluster-level covariates and model them using a dependent Bayesian nonparametric model. The resulting process is flexible and easy to fit using an existing R package. The application of the model to our motivating example showed that, contrary to intuition, those diagnosed during a period of time in the 1990s in more rural and less affluent Iowan counties survived breast cancer better. Additional analyses showed the opposite trend for earlier time windows. We conjecture that this anomaly has to be due to increased hormone replacement therapy treatments prescribed to more urban and affluent subpopulations.

stat.AP