SearcharxivSearch

arXiv subjects

Naoki Nonaka

Publications and source records attributed to Naoki Nonaka.

6 recordsLinked to original sources

Boosting ECG Classification Performance by Pre-training with Synthesized Data

Deep Neural Networks (DNNs) typically require extensive datasets for effective training. In the medical domain, acquiring large-scale data is often challenging due to privacy concerns and the rarity of certain diseases. To address this data scarcity, we investigate the efficacy of training DNN models using synthetic data, generated based on domain-specific medical knowledge. Specifically, we develop a knowledge-driven Gaussian-composition synthesis algorithm for single-lead II ECGs, in which each heartbeat is represented by Gaussian-shaped P, Q, R, S, and T wave components. Using this simulator, we generate synthetic data for four abnormal electrocardiogram (ECG) classes: atrial fibrillation (AF), atrial flutter (AFLT), premature ventricular complex (PVC), and Wolff-Parkinson-White Syndrome (WPW). We evaluate the utility of this synthetic data by conducting abnormal ECG classification using ten different DNN architectures. Our results demonstrate that synthetic-to-real training improves classification performance for three of the four target abnormalities, with the largest architecture-averaged gain of $33.2\%$ observed for AFLT. Further analysis reveals that the performance enhancement from synthetic data is more pronounced with smaller real-world datasets. These findings suggest that domain-knowledge-based synthetic ECGs can serve as a useful pre-training resource, particularly in scenarios where real-world data are limited or difficult to obtain.

cs.LG

Syn-STARTS: Synthesized START Triage Scenario Generation Framework for Scalable LLM Evaluation

Triage is a critically important decision-making process in mass casualty incidents (MCIs) to maximize victim survival rates. While the role of AI in such situations is gaining attention for making optimal decisions within limited resources and time, its development and performance evaluation require benchmark datasets of sufficient quantity and quality. However, MCIs occur infrequently, and sufficient records are difficult to accumulate at the scene, making it challenging to collect large-scale realworld data for research use. Therefore, we developed Syn-STARTS, a framework that uses LLMs to generate triage cases, and verified its effectiveness. The results showed that the triage cases generated by Syn-STARTS were qualitatively indistinguishable from the TRIAGE open dataset generated by manual curation from training materials. Furthermore, when evaluating the LLM accuracy using hundreds of cases each from the green, yellow, red, and black categories defined by the standard triage method START, the results were found to be highly stable. This strongly indicates the possibility of synthetic data in developing high-performance AI models for severe and critical medical situations.

cs.AI

Efficient HLA imputation from sequential SNPs data by Transformer

Human leukocyte antigen (HLA) genes are associated with a variety of diseases, however direct typing of HLA is time and cost consuming. Thus various imputation methods using sequential SNPs data have been proposed based on statistical or deep learning models, e.g. CNN-based model, named DEEP*HLA. However, imputation efficiency is not sufficient for in frequent alleles and a large size of reference panel is required. Here, we developed a Transformer-based model to impute HLA alleles, named "HLA Reliable IMputatioN by Transformer (HLARIMNT)" to take advantage of sequential nature of SNPs data. We validated the performance of HLARIMNT using two different reference panels; Pan-Asian reference panel (n = 530) and Type 1 Diabetes Genetics Consortium (T1DGC) reference panel (n = 5,225), as well as the mixture of those two panels (n = 1,060). HLARIMNT achieved higher accuracy than DEEP*HLA by several indices, especially for infrequent alleles. We also varied the size of data used for training, and HLARIMNT imputed more accurately among any size of training data. These results suggest that Transformer-based model may impute efficiently not only HLA types but also any other gene types from sequential SNPs data.

q-bio.GN

Data Augmentation for Electrocardiogram Classification with Deep Neural Network

Electrocardiogram (ECG) is the most crucial monitoring modality to diagnose cardiovascular events. Precise and automatic detection of abnormal ECG patterns is beneficial to both physicians and patients. In the automatic detection of abnormal ECG patterns, deep neural networks (DNNs) have shown significant achievements. However, DNNs require large amount of labeled data, which are often expensive to obtain. On the other hand, recent research have shown by randomly combining data augmentations can improve image classification accuracy. Thus, in this work we explore data augmentation suitable for ECG data and propose ECG Augment. We show by introducing ECG Augment, we can improve classification of atrial fibrillation with single lead ECG data, without changing an architecture of DNN.

eess.SP

General-to-Detailed GAN for Infrequent Class Medical Images

Deep learning has significant potential for medical imaging. However, since the incident rate of each disease varies widely, the frequency of classes in a medical image dataset is imbalanced, leading to poor accuracy for such infrequent classes. One possible solution is data augmentation of infrequent classes using synthesized images created by Generative Adversarial Networks (GANs), but conventional GANs also require certain amount of images to learn. To overcome this limitation, here we propose General-to-detailed GAN (GDGAN), serially connected two GANs, one for general labels and the other for detailed labels. GDGAN produced diverse medical images, and the network trained with an augmented dataset outperformed other networks using existing methods with respect to Area-Under-Curve (AUC) of Receiver Operating Characteristic (ROC) curve.

cs.CV

Observation of Diffuse Cosmic and Atmospheric Gamma Rays at Balloon Altitudes with an Electron-tracking Compton Camera

We observed diffuse cosmic and atmospheric gamma rays at balloon altitudes with the Sub-MeV gamma-ray Imaging Loaded-on-balloon Experiment I (SMILE-I) as the first step toward a future all-sky survey with a high sensitivity. SMILE-I employed an electron-tracking Compton camera comprised of a gaseous electron tracker as a Compton-scattering target and a scintillation camera as an absorber. The balloon carrying the SMILE-I detector was launched from the Sanriku Balloon Center of the Institute of Space and Astronomical Science/Japan Space Exploration Agency on September 1, 2006, and the flight lasted for 6.8 hr, including level flight for 4.1 hr at an altitude of 32-35 km. During the level flight, we successfully detected 420 downward gamma rays between 100 keV and 1 MeV at zenith angles below 60 degrees. To obtain the flux of diffuse cosmic gamma rays, we first simulated their scattering in the atmosphere using Geant4, and for gamma rays detected at an atmospheric depth of 7.0 g cm-2, we found that 50% and 21% of the gamma rays at energies of 150 keV and 1 MeV, respectively, were scattered in the atmosphere prior to reaching the detector. Moreover, by using Geant4 simulations and the QinetiQ atmospheric radiation model, we estimated that the detected events consisted of diffuse cosmic and atmospheric gamma rays (79%), secondary photons produced in the instrument through the interaction between cosmic rays and materials surrounding the detector (19%), and other particles (2%). The obtained growth curve was comparable to Ling's model, and the fluxes of diffuse cosmic and atmospheric gamma rays were consistent with the results of previous experiments. The expected detection sensitivity of a future SMILE experiment measuring gamma rays between 150 keV and 20 MeV was estimated from our SMILE-I results and was found to be ten times better than that of other experiments at around 1 MeV.

astro-ph.IM