SearcharxivSearch

arXiv subjects

Tatsuya Mikami

Publications and source records attributed to Tatsuya Mikami.

4 recordsLinked to original sources

Out-of-distribution Reject Option Method for Dataset Shift Problem in Early Disease Onset Prediction

Machine learning is increasingly used to predict lifestyle-related disease onset using health and medical data. However, its predictive accuracy for use is often hindered by dataset shift, which refers to discrepancies in data distribution between the training and testing datasets. This issue leads to the misclassification of out-of-distribution (OOD) data. To diminish dataset shift in real-world settings, this paper proposes the out-of-distribution reject option for prediction (ODROP). This method integrates an OOD detection model to preclude OOD data from the prediction phase. We used two real-world health checkup datasets (Hirosaki and Wakayama) with dataset shift, across three disease onset prediction tasks: diabetes, dyslipidemia, and hypertension. Both components of ODROP method -- the OOD detection model and the prediction model -- were trained on the Hirosaki dataset. We assessed the effectiveness of ODROP on the Wakayama dataset using AUROC-rejection rate curve plot. In the five OOD detection approaches (the variational autoencoder, neural network ensemble std, neural network ensemble epistemic, neural network energy, and neural network gaussian mixture based energy measurement), the variational autoencoder method demonstrated notably higher stability and a greater improvement in AUROC. For example, in the Wakayama dataset, the AUROC for diabetes onset increased from 0.80 without ODROP to 0.90 at a 31.1% rejection rate, and for dyslipidemia, it improved from 0.70 without ODROP to 0.76 at a 34% rejection rate. In addition, we categorized dataset shifts into two types using SHAP clustering -- those that considerably affect predictions and those that do not. This study is the first to apply OOD detection to actual health and medical data, demonstrating its potential to substantially improve the accuracy and reliability of disease prediction models amidst dataset shift.

cs.LG

Individual health-disease phase diagrams for disease prevention based on machine learning

Early disease detection and prevention methods based on effective interventions are gaining attention. Machine learning technology has enabled precise disease prediction by capturing individual differences in multivariate data. Progress in precision medicine has revealed that substantial heterogeneity exists in health data at the individual level and that complex health factors are involved in the development of chronic diseases. However, it remains a challenge to identify individual physiological state changes in cross-disease onset processes because of the complex relationships among multiple biomarkers. Here, we present the health-disease phase diagram (HDPD), which represents a personal health state by visualizing the boundary values of multiple biomarkers that fluctuate early in the disease progression process. In HDPDs, future onset predictions are represented by perturbing multiple biomarker values while accounting for dependencies among variables. We constructed HDPDs for 11 non-communicable diseases (NCDs) from a longitudinal health checkup cohort of 3,238 individuals, comprising 3,215 measurement items and genetic data. Improvement of biomarker values to the non-onset region in HDPD significantly prevented future disease onset in 7 out of 11 NCDs. Our results demonstrate that HDPDs can represent individual physiological states in the onset process and be used as intervention goals for disease prevention.

cs.LG

Covering monotonicity of the limit shapes of first passage percolation on crystal lattices

This paper studies the first passage percolation (FPP) model: each edge in the cubic lattice is assigned a random passage time, and consideration is given to the behavior of the percolation region $B(t)$, which consists of those vertices that can be reached from the origin within a time $t > 0$. Cox and Durrett showed the shape theorem for the percolation region, saying that the normalized region $B(t)/t$ converges to some limit shape $\mathcal{B}$. This paper introduces a general FPP model defined on crystal lattices, and shows the monotonicity of the limit shapes under covering maps, thereby providing insight into the limit shape of the cubic FPP model.

math.PR

Percolation on Homology Generators in Codimension One

This paper introduces a new percolation model motivated from polymer materials. The mathematical model is defined over a random cubical set in the $d$-dimensional space $\mathbb{R}^d$ and focuses on generations and percolations of $(d-1)$-dimensional holes as higher dimensional topological objects. Here, the random cubical set is constructed by the union of unit faces in dimension $d-1$ which appear randomly and independently with probability $p$, and holes are formulated by the homology generators. Under this model, the upper and lower estimates of the critical probability $p_c^{\rm hole}$ of the hole percolation are shown in this paper, implying the existence of the phase transition. The uniqueness of infinite hole cluster is also proven. This result shows that, when $p > p_c^{\rm hole}$, the probability $P_p(x^*\overset{\rm hole}{\longleftrightarrow} y^*)$ that two points in the dual lattice $(\mathbb{Z}^d)^*$ belong to the same hole cluster is uniformly greater than 0.

math.PR