SearcharxivSearch

arXiv · 1905.08603

Overview of image-to-image translation by use of deep neural networks: denoising, super-resolution, modality conversion, and reconstruction in medical imaging

Abstract

Since the advent of deep convolutional neural networks (DNNs), computer vision has seen an extremely rapid progress that has led to huge advances in medical imaging. This article does not aim to cover all aspects of the field but focuses on a particular topic, image-to-image translation. Although the topic may not sound familiar, it turns out that many seemingly irrelevant applications can be understood as instances of image-to-image translation. Such applications include (1) noise reduction, (2) super-resolution, (3) image synthesis, and (4) reconstruction. The same underlying principles and algorithms work for various tasks. Our aim is to introduce some of the key ideas on this topic from a uniform point of view. We introduce core ideas and jargon that are specific to image processing by use of DNNs. Having an intuitive grasp of the core ideas of and a knowledge of technical terms would be of great help to the reader for understanding the existing and future applications. Most of the recent applications which build on image-to-image translation are based on one of two fundamental architectures, called pix2pix and CycleGAN, depending on whether the available training data are paired or unpaired. We provide computer codes which implement these two architectures with various enhancements. Our codes are available online with use of the very permissive MIT license. We provide a hands-on tutorial for training a model for denoising based on our codes. We hope that this article, together with the codes, will provide both an overview and the details of the key algorithms, and that it will serve as a basis for the development of new applications.

Explore related subjects

Keep this discovery

BibTeXRIS

Shizuo Kaji, Satoshi Kida. 2019-05-21. Overview of image-to-image translation by use of deep neural networks: denoising, super-resolution, modality conversion, and reconstruction in medical imaging. https://arxiv.org/abs/1905.08603

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Phase-contrast micro-CT for intra-operative breast tumour margin assessment using a microfocus x-ray source and photon-counting detector

Objective: Intra-operative tumour margin assessment during breast-conserving surgery requires rapid, high-resolution imaging of excised tissue, allowing the surgical team to take appropriate action within a single operation. This study evaluates a custom propagation-based phase-contrast micro-computed tomography (micro-CT) system designed to meet these clinical constraints without specialised optical elements. Methods: The experimental setup pairs a microfocus x-ray source with a photon-counting detector in a cone-beam geometry. We explore how the spatial coherence of the source can provide propagation-based phase contrast -- with no additional specialised optical elements -- and balance this against maximising the x-ray flux of the cone-beam geometry. System performance was evaluated across two anode target materials and filtration configurations at various tube power settings. Imaging capabilities were validated using anthropomorphic breast tissue phantoms and a formalin-fixed paraffin-embedded (FFPE) breast tissue specimen, with reconstructions compared against gold-standard histology. Results: An unfiltered tungsten target operated at 40 kVp yielded optimal image quality. The optimised system achieved high-resolution CT reconstructions of a 5 cm diameter sample with an isotropic voxel size of 40.7 $\upmu\text{m}$ in a scan time of 12 minutes. Reconstructed volumes demonstrated strong visual correlation with corresponding histology slides. Conclusion: Combining a microfocus source with a photon-counting detector enables high-resolution, phase-contrast micro-CT within a clinically viable timeframe, demonstrating strong potential for intra-operative margin assessment.

physics.med-ph

Understanding Search and Decision Errors in Liver Metastasis Detection and the Effects of Lower Radiation Dose

The detection performance of liver metastases decreases with the reduction of radiation dose, but misses are heterogeneous. Previous eye tracking work has characterized missed metastases into two categories: search errors i.e., the eyes never land on the lesion, and decision errors i.e., the lesion is seen but not recognized as malignant. We integrated three prior reader studies to answer this question. In all studies, radiologists interpreted the same set of 40 contrast enhanced abdominal CT exams containing 91 liver metastases whose locations had been previously marked. In two studies, the workstation recorded their gaze and eye movements. Using eye dwell times, metastases were classified as search-error-dominant (majority of misses had <2 sec gaze time) or decision-error-dominant (>2 sec gaze time). In the third study, exams were interpreted both at 120 and 200 quality reference mAs (QRM) by ten radiologists. The third study did not include eye tracking. Out of 91 liver metastases, we excluded 16 that were never missed in the eye tracking studies and used 75 liver metastases for the present study.

physics.med-ph

Develop and Optimize 5DCT Imaging Simulation and Reconstruction Methods

Purpose: To develop and optimize a 5DCT (3D + cardiac phase + respiratory phase) imaging simulation and reconstruction pipeline, and to compare two sinogram-space interpolation methods for reconstructing images at arbitrary combinations of cardiac and respiratory phase. Methods: Helical CT projections were simulated from the 4D XCAT phantom across a range of cardiac and respiratory motion states, with Poisson and electronic noise added. Ground-truth-matched volumes were generated at 5 cardiac phases and 10 respiratory amplitudes (50 total phase combinations). Because acquired projections are sparsely and unevenly distributed across this joint phase space, each target slice was reconstructed by interpolating rebinned sinogram rows to the target cardiac phase and respiratory amplitude, using either 2D scattered barycentric interpolation or 2D scattered local linear interpolation with a circular kernel for cardiac phase. Reconstructed volumes were compared to phantom ground truth using mean absolute error (MAE), and to conventional respiratory-gated 4DCT (r4DCT) reconstructed from the same simulated data. Results: Both interpolation methods eliminated the severe axial misalignment artifacts present when helical projections were reconstructed without phase-space interpolation. Local linear interpolation achieved lower MAE than barycentric interpolation across most tested conditions, with the largest improvement at low pitch. The 5DCT pipeline also produced respiratory-only volumes with fewer residual cardiac-motion artifacts than conventional r4DCT reconstructed from the same projection data, including at standard clinical pitch (0.1). Conclusions: 5DCT reconstruction using sinogram-space interpolation is feasible and can jointly resolve cardiac and respiratory motion with better accuracy than conventional 4DCT reconstruction.

physics.med-ph