SearcharxivSearch

arXiv subjects

Daniel Zelterman

Publications and source records attributed to Daniel Zelterman.

7 recordsLinked to original sources

Bayesian local exchangeability design for phase II basket trials

We propose an information borrowing strategy for the design and monitoring of phase II basket trials based on the local multisource exchangeability assumption between baskets (disease types). In our proposed local-MEM framework, information borrowing is only allowed to occur locally, i.e., among baskets with similar response rate and the amount of information borrowing is determined by the level of similarity in response rate, whereas baskets not considered similar are not allowed to share information. We construct a two-stage design for phase II basket trials using the proposed strategy. The proposed method is compared to competing Bayesian methods and Simon's two-stage design in a variety of simulation scenarios. We demonstrate the proposed method is able to maintain the family-wise type I error rate at a reasonable level and has desirable basket-wise power compared to Simon's two-stage design. In addition, our method is computationally efficient compared to existing Bayesian methods in that the posterior profiles of interest can be derived explicitly without the need for sampling algorithms.

stat.ME

Distributions associated with simultaneous multiple hypothesis testing

We develop the distribution of the number of hypotheses found to be statistically significant using the rule from Benjamini and Hochberg (1995) for controlling the false discovery rate (FDR). This distribution has both a small sample form and an asymptotic expression for testing many independent hypotheses simultaneously. We propose a parametric distribution $\,Ψ_I(\cdot)\,$ to approximate the marginal distribution of p-values under a non-uniform alternative hypothesis. This distribution is useful when there are many different alternative hypotheses and these are not individually well understood. We fit $\,Ψ_I\,$ to data from three cancer studies and use it to illustrate the distribution of the number of notable hypotheses observed in these examples. We model dependence of sampled p-values using a copula model and a latent variable approach. These methods can be combined to illustrate a power analysis in planning a large study on the basis of a smaller pilot study. We show the number of statistically significant p-values behaves approximately as a mixture of a normal and the Borel-Tanner distribution.

stat.ME

The maximum negative hypergeometric distribution

An urn contains a known number of balls of two different colors. We describe the random variable counting the smallest number of draws needed in order to observe at least $\,c\,$ of both colors when sampling without replacement for a pre-specified value of $\,c=1,2,\ldots\,$. This distribution is the finite sample analogy to the maximum negative binomial distribution described by Zhang, Burtness, and Zelterman (2000). We describe the modes, approximating distributions, and estimation of the contents of the urn.

math.ST

A Stopped Negative Binomial Distribution

This paper introduces a new discrete distribution suggested by curtailed sampling rules common in early-stage clinical trials. We derive the distribution of the smallest number of independent Bernoulli(p) trials needed in order to observe either s successes or t failures. The closed form expression for the distribution as well as the compound distribution are derived. Properties of the distribution are shown and discussed. A case study is presented showing how the distribution can be used to monitor sequential enrollment of clinical trials with binary outcomes as well as providing post-hoc analysis of completed trials.

math.ST

Markov counting models for correlated binary responses

We propose a class of continuous-time Markov counting processes for analyzing correlated binary data and establish a correspondence between these models and sums of exchangeable Bernoulli random variables. Our approach generalizes many previous models for correlated outcomes, admits easily interpretable parameterizations, allows different cluster sizes, and incorporates ascertainment bias in a natural way. We demonstrate several new models for dependent outcomes and provide algorithms for computing maximum likelihood estimates. We show how to incorporate cluster-specific covariates in a regression setting and demonstrate improved fits to well-known datasets from familial disease epidemiology and developmental toxicology.

stat.ME

A Spatial Simulation Approach to Account for Protein Structure When Identifying Non-Random Somatic Mutations

Background: Current research suggests that a small set of "driver" mutations are responsible for tumorigenesis while a larger body of "passenger" mutations occurs in the tumor but does not progress the disease. Due to recent pharmacological successes in treating cancers caused by driver mutations, a variety of of methodologies that attempt to identify such mutations have been developed. Based on the hypothesis that driver mutations tend to cluster in key regions of the protein, the development of cluster identification algorithms has become critical. Results: We have developed a novel methodology, SpacePAC (Spatial Protein Amino acid Clustering), that identifies mutational clustering by considering the protein tertiary structure directly in 3D space. By combining the mutational data in the Catalogue of Somatic Mutations in Cancer (COSMIC) and the spatial information in the Protein Data Bank (PDB), SpacePAC is able to identify novel mutation clusters in many proteins such as FGFR3 and CHRM2. In addition, SpacePAC is better able to localize the most significant mutational hotspots as demonstrated in the cases of BRAF and ALK. The R package is available on Bioconductor at: http://www.bioconductor.org/packages/release/bioc/html/SpacePAC.html Conclusion: SpacePAC adds a valuable tool to the identification of mutational clusters while considering protein tertiary structure

q-bio.GN

A Two-Stage, Phase II Clinical Trial Design with Nested Criteria for Early Stopping and Efficacy

We propose a two-stage design for a clinical trial with an early stopping rule for safety. We use different criteria to assess early stopping and efficacy. The early stopping rule is based on a criteria that can be determined more quickly than that of efficacy. These separate criteria are also nested in the sense that efficacy is a special case of, but not identical to, the early stopping criteria. The design readily allows for planning in terms of statistical significance, power, and expected sample size necessary to assess an early stopping rule. This method is illustrated with a Phase II design comparing patients treated for lung cancer with a novel drug combination to those treated using historical control. In this example, the early stopping rule is based on the numbers of patients who exhibit progression-free survival (PFS) at 2 months post treatment follow-up and efficacy is judged by the number of patients who have PFS at 6 months.

stat.ME