SearcharxivSearch

arXiv subjects

J. Jack Lee

Publications and source records attributed to J. Jack Lee.

5 recordsLinked to original sources

Knowledge Distillation Decision Tree for Unravelling Black-box Machine Learning Models

Machine learning models, particularly the black-box models, are widely favored for their outstanding predictive capabilities. However, they often face scrutiny and criticism due to the lack of interpretability. Paradoxically, their strong predictive capabilities may indicate a deep understanding of the underlying data, implying significant potential for interpretation. Leveraging the emerging concept of knowledge distillation, we introduce the method of knowledge distillation decision tree (KDDT). This method enables the distillation of knowledge about the data from a black-box model into a decision tree, thereby facilitating the interpretation of the black-box model. Essential attributes for a good interpretable model include simplicity, stability, and predictivity. The primary challenge of constructing interpretable tree lies in ensuring structural stability under the randomness of the training data. KDDT is developed with the theoretical foundations demonstrating that structure stability can be achieved under mild assumptions. Furthermore, we propose the hybrid KDDT to achieve both simplicity and predictivity. An efficient algorithm is provided for constructing the hybrid KDDT. Simulation studies and a real-data analysis validate the hybrid KDDT's capability to deliver accurate and reliable interpretations. KDDT is an excellent interpretable model with great potential for practical applications.

stat.ME

Bayesian Clustering Prior with Overlapping Indices for Effective Use of Multisource External Data

The use of external data in clinical trials offers numerous advantages, such as reducing the number of patients, increasing study power, and shortening trial durations. In Bayesian inference, information in external data can be transferred into an informative prior for future borrowing (i.e., prior synthesis). However, multisource external data often exhibits heterogeneity, which can lead to information distortion during the prior synthesis. Clustering helps identifying the heterogeneity, enhancing the congruence between synthesized prior and external data, thereby preventing information distortion. Obtaining optimal clustering is challenging due to the trade-off between congruence with external data and robustness to future data. We introduce two overlapping indices: the overlapping clustering index (OCI) and the overlapping evidence index (OEI). Using these indices alongside a K-Means algorithm, the optimal clustering of external data can be identified by balancing the trade-off. Based on the clustering result, we propose a prior synthesis framework to effectively borrow information from multisource external data. By incorporating the (robust) meta-analytic predictive prior into this framework, we develop (robust) Bayesian clustering MAP priors. Simulation studies and real-data analysis demonstrate their superiority over commonly used priors in the presence of heterogeneity. Since the Bayesian clustering priors are constructed without needing data from the prospective study to be conducted, they can be applied to both study design and data analysis in clinical trials or experiments.

stat.ME

BOP2-TE: Bayesian Optimal Phase 2 Design for Jointly Monitoring Efficacy and Toxicity with Application to Dose Optimization

We propose a Bayesian optimal phase 2 design for jointly monitoring efficacy and toxicity, referred to as BOP2-TE, to improve the operating characteristics of the BOP2 design proposed by Zhou et al. (2017). BOP2-TE utilizes a Dirichlet-multinomial model to jointly model the distribution of toxicity and efficacy endpoints, making go/no-go decisions based on the posterior probability of toxicity and futility. In comparison to the original BOP2 and other existing designs, BOP2-TE offers the advantage of providing rigorous type I error control in cases where the treatment is toxic and futile, effective but toxic, or safe but futile, while optimizing power when the treatment is effective and safe. As a result, BOP2-TE enhances trial safety and efficacy. We also explore the incorporation of BOP2-TE into multiple-dose randomized trials for dose optimization, and consider a seamless design that integrates phase I dose finding with phase II randomized dose optimization. BOP2-TE is user-friendly, as its decision boundary can be determined prior to the trial's onset. Simulations demonstrate that BOP2-TE possesses desirable operating characteristics. We have developed a user-friendly web application as part of the BOP2 app, which is freely available at www.trialdesign.org.

stat.ME

Overlapping Indices for Dynamic Information Borrowing in Bayesian Hierarchical Modeling

Bayesian hierarchical model (BHM) has been widely used in synthesizing information across subgroups. Identifying heterogeneity in the data and determining proper strength of borrow have long been central goals pursued by researchers. Because these two goals are interconnected, we must consider them together. This joint consideration presents two fundamental challenges: (1) How can we balance the trade-off between homogeneity within the cluster and information gain through borrowing? (2) How can we determine the borrowing strength dynamically in different clusters? To tackle challenges, first, we develop a theoretical framework for heterogeneity identification and dynamic information borrowing in BHM. Then, we propose two novel overlapping indices: the overlapping clustering index (OCI) for identifying the optimal clustering result and the overlapping borrowing index (OBI) for assigning proper borrowing strength to clusters. By incorporating these indices, we develop a new method BHMOI (Bayesian hierarchical model with overlapping indices). BHMOI includes a novel weighted K-Means clustering algorithm by maximizing OCI to obtain optimal clustering results, and embedding OBI into BHM for dynamically borrowing within clusters. BHMOI can achieve efficient and robust information borrowing with desirable properties. Examples and simulation studies are provided to demonstrate the effectiveness of BHMOI in heterogeneity identification and dynamic information borrowing.

stat.ME

Incorporating historical information to improve phase I clinical trial designs

Incorporating historical data or real-world evidence has a great potential to improve the efficiency of phase I clinical trials and to accelerate drug development. For model-based designs, such as the continuous reassessment method (CRM), this can be conveniently carried out by specifying a "skeleton," i.e., the prior estimate of dose limiting toxicity (DLT) probability at each dose. In contrast, little work has been done to incorporate historical data or real-world evidence into model-assisted designs, such as the Bayesian optimal interval (BOIN), keyboard, and modified toxicity probability interval (mTPI) designs. This has led to the misconception that model-assisted designs cannot incorporate prior information. In this paper, we propose a unified framework that allows for incorporating historical data or real-world evidence into model-assisted designs. The proposed approach uses the well-established "skeleton" approach, combined with the concept of prior effective sample size, thus it is easy to understand and use. More importantly, our approach maintains the hallmark of model-assisted designs: simplicity---the dose escalation/de-escalation rule can be tabulated prior to the trial conduct. Extensive simulation studies show that the proposed method can effectively incorporate prior information to improve the operating characteristics of model-assisted designs, similarly to model-based designs.

stat.ME