SearcharxivSearch

arXiv subjects

Ziping Liu

Publications and source records attributed to Ziping Liu.

9 recordsLinked to original sources

NGSE-Corr: A technique for objective clinical evaluation of quantitative-imaging methods without a gold standard

Objective evaluation of quantitative-imaging (QI) methods based on how reliably they measure true values is important for clinical translation. Performing such evaluation with patient data is highly desirable but hindered by the lack of gold standards. To address this challenge, advancing on previous studies, we propose a no-gold-standard evaluation technique, NGSE-Corr, that objectively evaluates QI methods without true values. The technique assumes a linear stochastic relationship between true and measured values, characterized by a slope, bias, and multivariate Gaussian-distributed noise term that models correlated noise across QI methods. We derive a maximum-likelihood approach to estimate these parameters using only measured values. From the estimates, we compute noise-to-slope ratio (NSR) to rank QI methods based on precision. Numerical experiments showed that NGSE-Corr reliably estimated the NSR, accurately ranked methods, and maintained performance even when assumptions made by the technique were partially violated. We also validated NGSE-Corr in an in silico imaging trial to rank three quantitative SPECT methods for measuring regional activity uptake in patients with bone metastatic castrate-resistant prostate cancer treated with radium-223. NGSE-Corr correctly identified the most precise QI method and ranked the methods for 95% (95% CI, 89%-98%) and 91% (95% CI, 84%-95%) of trials, respectively, with data from 50 patients. Performance further improved with larger cohorts. With 200 patients, NGSE-Corr yielded same rankings as those obtained with true values across all trial instances. These findings demonstrate the ability of NGSE-Corr to accurately rank QI methods without gold standards and motivate clinical validation and broader applications.

physics.med-ph

Observer study-based evaluation of TGAN architecture used to generate oncological PET images

The application of computer-vision algorithms in medical imaging has increased rapidly in recent years. However, algorithm training is challenging due to limited sample sizes, lack of labeled samples, as well as privacy concerns regarding data sharing. To address these issues, we previously developed (Bergen et al. 2022) a synthetic PET dataset for Head and Neck (H and N) cancer using the temporal generative adversarial network (TGAN) architecture and evaluated its performance segmenting lesions and identifying radiomics features in synthesized images. In this work, a two-alternative forced-choice (2AFC) observer study was performed to quantitatively evaluate the ability of human observers to distinguish between real and synthesized oncological PET images. In the study eight trained readers, including two board-certified nuclear medicine physicians, read 170 real/synthetic image pairs presented as 2D-transaxial using a dedicated web app. For each image pair, the observer was asked to identify the real image and input their confidence level with a 5-point Likert scale. P-values were computed using the binomial test and Wilcoxon signed-rank test. A heat map was used to compare the response accuracy distribution for the signed-rank test. Response accuracy for all observers ranged from 36.2% [27.9-44.4] to 63.1% [54.8-71.3]. Six out of eight observers did not identify the real image with statistical significance, indicating that the synthetic dataset was reasonably representative of oncological PET images. Overall, this study adds validity to the realism of our simulated H&N cancer dataset, which may be implemented in the future to train AI algorithms while favoring patient confidentiality and privacy protection.

physics.med-ph

Need for objective task-based evaluation of AI-based segmentation methods for quantitative PET

Artificial intelligence (AI)-based methods are showing substantial promise in segmenting oncologic positron emission tomography (PET) images. For clinical translation of these methods, assessing their performance on clinically relevant tasks is important. However, these methods are typically evaluated using metrics that may not correlate with the task performance. One such widely used metric is the Dice score, a figure of merit that measures the spatial overlap between the estimated segmentation and a reference standard (e.g., manual segmentation). In this work, we investigated whether evaluating AI-based segmentation methods using Dice scores yields a similar interpretation as evaluation on the clinical tasks of quantifying metabolic tumor volume (MTV) and total lesion glycolysis (TLG) of primary tumor from PET images of patients with non-small cell lung cancer. The investigation was conducted via a retrospective analysis with the ECOG-ACRIN 6668/RTOG 0235 multi-center clinical trial data. Specifically, we evaluated different structures of a commonly used AI-based segmentation method using both Dice scores and the accuracy in quantifying MTV/TLG. Our results show that evaluation using Dice scores can lead to findings that are inconsistent with evaluation using the task-based figure of merit. Thus, our study motivates the need for objective task-based evaluation of AI-based segmentation methods for quantitative PET.

physics.med-ph

A tissue-fraction estimation-based segmentation method for quantitative dopamine transporter SPECT

Quantitative measures of dopamine transporter (DaT) uptake in caudate, putamen, and globus pallidus (GP) have potential as biomarkers for measuring the severity of Parkinson disease. Reliable quantification of this uptake requires accurate segmentation of the considered regions. However, segmentation of these regions from DaT-SPECT images is challenging, a major reason being partial-volume effects (PVEs), which arise from the limited system resolution and reconstruction of images over finite-sized voxel grids. The latter leads to tissue-fraction effects (TFEs). Thus, there is an important need for methods that can account for the PVEs, including the TFEs, and accurately segment DaT-SPECT images. The purpose of this study is to design and objectively evaluate a fully automated tissue-fraction estimation-based segmentation method that segments the caudate, putamen, and GP from DaT-SPECT images. The proposed method estimates the posterior mean of the fractional volumes occupied by the caudate, putamen, and GP within each voxel of a 3-D DaT-SPECT image. The estimate is obtained by minimizing a cost function based on the binary cross-entropy loss between the true and estimated fractional volumes over a population of SPECT images. Evaluations using clinically guided highly realistic simulation studies show that the proposed method accurately segmented the caudate, putamen, and GP with high mean Dice similarity coefficients ~ 0.80 and significantly outperformed (p < 0.01) all other considered segmentation methods. Further, objective evaluation of the proposed method on the task of quantifying regional uptake shows that the method yielded reliable quantification with low ensemble normalized root mean square error (NRMSE) < 20% for all the considered regions. The results motivate further evaluation of the method with physical-phantom and patient studies.

physics.med-ph

A Bayesian approach to tissue-fraction estimation for oncological PET segmentation

Tumor segmentation in oncological PET is challenging, a major reason being the partial-volume effects that arise due to low system resolution and finite voxel size. The latter results in tissue-fraction effects, i.e. voxels contain a mixture of tissue classes. Conventional segmentation methods are typically designed to assign each voxel in the image as belonging to a certain tissue class. Thus, these methods are inherently limited in modeling tissue-fraction effects. To address the challenge of accounting for partial-volume effects, and in particular, tissue-fraction effects, we propose a Bayesian approach to tissue-fraction estimation for oncological PET segmentation. Specifically, this Bayesian approach estimates the posterior mean of fractional volume that the tumor occupies within each voxel of the image. The proposed method, implemented using a deep-learning-based technique, was first evaluated using clinically realistic 2-D simulation studies with known ground truth, in the context of segmenting the primary tumor in PET images of patients with lung cancer. The evaluation studies demonstrated that the method accurately estimated the tumor-fraction areas and significantly outperformed widely used conventional PET segmentation methods, including a U-net-based method, on the task of segmenting the tumor. In addition, the proposed method was relatively insensitive to partial-volume effects and yielded reliable tumor segmentation for different clinical-scanner configurations. The method was then evaluated using clinical images of patients with stage IIB/III non-small cell lung cancer from ACRIN 6668/RTOG 0235 multi-center clinical trial. Here, the results showed that the proposed method significantly outperformed all other considered methods and yielded accurate tumor segmentation on patient images with Dice similarity coefficient (DSC) of 0.82 (95 % CI: [0.78, 0.86]).

physics.med-ph

No-gold-standard evaluation of quantitative imaging methods in the presence of correlated noise

Objective evaluation of quantitative imaging (QI) methods with patient data is highly desirable, but is hindered by the lack or unreliability of an available gold standard. To address this issue, techniques that can evaluate QI methods without access to a gold standard are being actively developed. These techniques assume that the true and measured values are linearly related by a slope, bias, and Gaussian-distributed noise term, where the noise between measurements made by different methods is independent of each other. However, this noise arises in the process of measuring the same quantitative value, and thus can be correlated. To address this limitation, we propose a no-gold-standard evaluation (NGSE) technique that models this correlated noise by a multi-variate Gaussian distribution parameterized by a covariance matrix. We derive a maximum-likelihood-based approach to estimate the parameters that describe the relationship between the true and measured values, without any knowledge of the true values. We then use the estimated slopes and diagonal elements of the covariance matrix to compute the noise-to-slope ratio (NSR) to rank the QI methods on the basis of precision. The proposed NGSE technique was evaluated with multiple numerical experiments. Our results showed that the technique reliably estimated the NSR values and yielded accurate rankings of the considered methods for ~ 83% of 160 trials. In particular, the technique correctly identified the most precise method for ~ 97% of the trials. Overall, this study demonstrates the efficacy of the NGSE technique to accurately rank different QI methods when the correlated noise is present, and without access to any knowledge of the ground truth. The results motivate further validation of this technique with realistic simulation studies and patient data.

physics.med-ph

Objective task-based evaluation of artificial intelligence-based medical imaging methods: Framework, strategies and role of the physician

Artificial intelligence (AI)-based methods are showing promise in multiple medical-imaging applications. Thus, there is substantial interest in clinical translation of these methods, requiring in turn, that they be evaluated rigorously. In this paper, our goal is to lay out a framework for objective task-based evaluation of AI methods. We will also provide a list of tools available in the literature to conduct this evaluation. Further, we outline the important role of physicians in conducting these evaluation studies. The examples in this paper will be proposed in the context of PET with a focus on neural-network-based methods. However, the framework is also applicable to evaluate other medical-imaging modalities and other types of AI methods.

physics.med-ph

Observer study-based evaluation of a stochastic and physics-based method to generate oncological PET images

Objective evaluation of new and improved methods for PET imaging requires access to images with ground truth, as can be obtained through simulation studies. However, for these studies to be clinically relevant, it is important that the simulated images are clinically realistic. In this study, we develop a stochastic and physics-based method to generate realistic oncological two-dimensional (2-D) PET images, where the ground-truth tumor properties are known. The developed method extends upon a previously proposed approach. The approach captures the observed variabilities in tumor properties from actual patient population. Further, we extend that approach to model intra-tumor heterogeneity using a lumpy object model. To quantitatively evaluate the clinical realism of the simulated images, we conducted a human-observer study. This was a two-alternative forced-choice (2AFC) study with trained readers (five PET physicians and one PET physicist). Our results showed that the readers had an average of ~ 50% accuracy in the 2AFC study. Further, the developed simulation method was able to generate wide varieties of clinically observed tumor types. These results provide evidence for the application of this method to 2-D PET imaging applications, and motivate development of this method to generate 3-D PET images.

physics.med-ph

A no-gold-standard technique to objectively evaluate quantitative imaging methods using patient data: Theory

Objective evaluation of quantitative imaging (QI) methods using measurements directly obtained from patient images is highly desirable but hindered by the non-availability of gold standards. To address this issue, statistical techniques have been proposed to objectively evaluate QI methods without a gold standard. These techniques assume that the measured and true values are linearly related by a slope, bias, and normally distributed noise term, where it is assumed that the noise term between the different methods is independent. However, the noise could be correlated since it arises in the process of measuring the same true value. To address this issue, we propose a new no-gold-standard evaluation (NGSE) technique that models this noise as a multivariate normally distributed term, characterized by a covariance matrix. In this manuscript, we derive a maximum-likelihood-based technique that, without any knowledge of the true QI values, estimates the slope, bias, and covariance matrix terms. These are then used to rank the methods on the basis of precision of the measured QI values. Overall, the derivation demonstrates the mathematical premise behind the proposed NGSE technique.

stat.ME