SearcharxivSearch

arXiv subjects

Guillaume Blanc

Publications and source records attributed to Guillaume Blanc.

14 recordsLinked to original sources

Explainable Graph-theoretical Machine Learning with Application to Alzheimer's Disease Prediction

Dementia affects over 55 million people worldwide, projected to reach 139 million by 2050, with Alzheimer's disease (AD) accounting for 60-70% of cases. AD is associated with disruptions in metabolic brain connectivity. Detecting these disruptions early is crucial for AD management. FDG-PET is a useful tool for identifying such impairments. However, most studies rely on group-level analyses or thresholding, potentially masking individual differences and overlooking weaker yet biologically critical brain connections. Moreover, AD prediction largely focuses on univariate rather than multivariate outcomes. To address this, we introduce explainable graph-theoretical machine learning (XGML), a framework for constructing individual metabolic brain graphs and identifying subgraphs most predictive of multivariate disease-related outcomes. Using Alzheimer's Disease Neuroimaging Initiative (ADNI) FDG-PET data, we compared six graph representations against three non-graph baselines, each with six machine learning models using repeated stratified 3-fold cross-validation (10 repeats). The best configuration combined kernel density estimation with Hellinger distance and random forest. Across eight cognitive scores, it reached an overall Fisher-z-averaged Pearson correlation of r=0.595, with strongest performance for ADAS13 (r=0.67), ADAS11 (r=0.65), and ADASQ4 (r=0.62). We identified key edges that were jointly but differentially predictive across outcomes, suggesting their potential as network biomarkers of cognitive decline. Preliminary external feasibility validation on an OASIS3 cohort yielded weak predictive performance for CDRSB (r=0.26) and MMSE (r=0.18), likely reflecting cohort, protocol, and diagnostic differences. Overall, our results suggest the promise of graph-theoretical machine learning for biomarker discovery, disease prediction, and understanding the neural mechanisms underlying AD.

cs.LG

Random burning of the Euclidean lattice

The burning number of a graph is the minimal number of steps that are needed to burn all of its vertices, with the following burning procedure: at each step, one can choose a point to set on fire, and the fire propagates constantly at unit speed along the edges of the graph. In this paper, we consider two natural random burning procedures in the discrete Euclidean torus $\mathbb{T}_n^d$, in which the points that we set on fire at each step are random variables. Our main result deals with the case where at each step, the law of the new point that we set on fire conditionally on the past is the uniform distribution on the complement of the set of vertices burned by the previous points. In this case, we prove that as $n\to\infty$, the corresponding random burning number (i.e, the first step at which the whole torus is burned) is asymptotic to $T\cdot n^{d/(d+1)}$ in probability, where $T=T(d)\in(0,\infty)$ is the explosion time of a so-called generalised Blasius equation.

math.PR

OPTIMUS: Predicting Multivariate Outcomes in Alzheimer's Disease Using Multi-modal Data amidst Missing Values

Alzheimer's disease, a neurodegenerative disorder, is associated with neural, genetic, and proteomic factors while affecting multiple cognitive and behavioral faculties. Traditional AD prediction largely focuses on univariate disease outcomes, such as disease stages and severity. Multimodal data encode broader disease information than a single modality and may, therefore, improve disease prediction; but they often contain missing values. Recent "deeper" machine learning approaches show promise in improving prediction accuracy, yet the biological relevance of these models needs to be further charted. Integrating missing data analysis, predictive modeling, multimodal data analysis, and explainable AI, we propose OPTIMUS, a predictive, modular, and explainable machine learning framework, to unveil the many-to-many predictive pathways between multimodal input data and multivariate disease outcomes amidst missing values. OPTIMUS first applies modality-specific imputation to uncover data from each modality while optimizing overall prediction accuracy. It then maps multimodal biomarkers to multivariate outcomes using machine-learning and extracts biomarkers respectively predictive of each outcome. Finally, OPTIMUS incorporates XAI to explain the identified multimodal biomarkers. Using data from 346 cognitively normal subjects, 608 persons with mild cognitive impairment, and 251 AD patients, OPTIMUS identifies neural and transcriptomic signatures that jointly but differentially predict multivariate outcomes related to executive function, language, memory, and visuospatial function. Our work demonstrates the potential of building a predictive and biologically explainable machine-learning framework to uncover multimodal biomarkers that capture disease profiles across varying cognitive landscapes. The results improve our understanding of the complex many-to-many pathways in AD.

cs.LG

Blow-up rate of solution to generalised Blasius equation

We identify the blow-up rate of a solution to a generalised Blasius equation, that we came across while studying a probabilistic model of "Poissonian burning" in Euclidean space. Our proof involves the study of the long-time behaviour of solutions to a Lotka--Volterra system.

math.CA

VisTA: Vision-Text Alignment Model with Contrastive Learning using Multimodal Data for Evidence-Driven, Reliable, and Explainable Alzheimer's Disease Diagnosis

Objective: Assessing Alzheimer's disease (AD) using high-dimensional radiology images is clinically important but challenging. Although Artificial Intelligence (AI) has advanced AD diagnosis, it remains unclear how to design AI models embracing predictability and explainability. Here, we propose VisTA, a multimodal language-vision model assisted by contrastive learning, to optimize disease prediction and evidence-based, interpretable explanations for clinical decision-making. Methods: We developed VisTA (Vision-Text Alignment Model) for AD diagnosis. Architecturally, we built VisTA from BiomedCLIP and fine-tuned it using contrastive learning to align images with verified abnormalities and their descriptions. To train VisTA, we used a constructed reference dataset containing images, abnormality types, and descriptions verified by medical experts. VisTA produces four outputs: predicted abnormality type, similarity to reference cases, evidence-driven explanation, and final AD diagnoses. To illustrate VisTA's efficacy, we reported accuracy metrics for abnormality retrieval and dementia prediction. To demonstrate VisTA's explainability, we compared its explanations with human experts' explanations. Results: Compared to 15 million images used for baseline pretraining, VisTA only used 170 samples for fine-tuning and obtained significant improvement in abnormality retrieval and dementia prediction. For abnormality retrieval, VisTA reached 74% accuracy and an AUC of 0.87 (26% and 0.74, respectively, from baseline models). For dementia prediction, VisTA achieved 88% accuracy and an AUC of 0.82 (30% and 0.57, respectively, from baseline models). The generated explanations agreed strongly with human experts' and provided insights into the diagnostic process. Taken together, VisTA optimize prediction, clinical reasoning, and explanation.

cs.CV

Geodesics in planar Poisson roads random metric

We study the structure of geodesics in the fractal random metric constructed by Kendall from a self-similar Poisson process of roads (i.e, lines with speed limits) in $\mathbb{R}^2$. In particular, we prove a conjecture of Kendall stating that geodesics do not pause en route, i.e, use roads of arbitrary small speed except at their endpoints. It follows that the geodesic frame of $\left(\mathbb{R}^2,T\right)$ is the set of points on roads. We also consider geodesic stars and hubs, and give a complete description of the local structure of geodesics around points on roads. Notably, we prove that leaving a road by driving off-road is never geodesic.

math.PR

The distance problem on measured metric spaces

What distributions arise as the distribution of the distance between two typical points in some measured metric space? This seems to be a surprisingly subtle problem. We conjecture that every distribution with a density function whose support contains $0$ does arise in this way, and give some partial results in that direction.

math.PR

Statistical Quantile Learning for Large, Nonlinear, and Additive Latent Variable Models

The studies of large-scale, high-dimensional data in fields such as genomics and neuroscience have injected new insights into science. Yet, despite advances, they are confronting several challenges, often simultaneously: lack of interpretability, nonlinearity, slow computation, inconsistency and uncertain convergence, and small sample sizes compared to high feature dimensions. Here, we propose a relatively simple, scalable, and consistent nonlinear dimension reduction method that can potentially address these issues in unsupervised settings. We call this method Statistical Quantile Learning (SQL) because, methodologically, it leverages on a quantile approximation of the latent variables together with standard nonparametric techniques (sieve or penalyzed methods). We show that estimating the model simplifies into a convex assignment matching problem; we derive its asymptotic properties; we show that the model is identifiable under few conditions. Compared to its linear competitors, SQL explains more variance, yields better separation and explanation, and delivers more accurate outcome prediction. Compared to its nonlinear competitors, SQL shows considerable advantage in interpretability, ease of use and computations in large-dimensional settings. Finally, we apply SQL to high-dimensional gene expression data (consisting of 20,263 genes from 801 subjects), where the proposed method identified latent factors predictive of five cancer types. The SQL package is available at https://github.com/jbodelet/SQL.

stat.ME

Fractal properties of the frontier in Poissonian coloring

We study a model of random partitioning by nearest-neighbor coloring from Poisson rain, introduced independently by Aldous and Preater. Given two initial points in $[0,1]^d$ respectively colored in red and blue, we let independent uniformly random points fall in $[0,1]^d$, and upon arrival, each point takes the color of the nearest point fallen so far. We prove that the colored regions converge in the Hausdorff sense towards two random closed subsets whose intersection, the frontier, has Hausdorff dimension strictly between $d-1$ and $d$, thus answering a conjecture raised by Aldous. However, several topological properties of the frontier remain elusive.

math.PR

Percolation with invariant Poisson processes of lines in the $3$-regular tree

In this paper, we study invariant Poisson processes of lines (i.e, bi-infinite geodesics) in the $3$-regular tree. More precisely, there exists a unique (up to multiplicative constant) locally finite Borel measure on the space of lines that is invariant under graph automorphisms, and we consider two Poissonian ways of playing with this invariant measure. First, following Benjamini, Jonasson, Schramm and Tykesson, we consider an invariant Poisson process of lines, and show that there is a critical value of the intensity below which a.s. the vacant set of the process percolates, and above which all its connected components are finite. Then, we consider an invariant Poisson process of roads (i.e, lines with speed limits), and show that there is a critical value of the parameter governing the speed limits of the roads below which a.s. one can drive to infinity in finite time using the road network generated by the process, and above which this is impossible.

math.PR

Fractal properties of Aldous-Kendall random metric

Investigating a model of scale-invariant random spatial network suggested by Aldous, Kendall constructed a random metric $T$ on $\mathbb{R}^d$, for which the distance between points is given by the optimal connection time, when travelling on the road network generated by a Poisson process of lines with a speed limit. In this paper, we look into some fractal properties of that random metric. In particular, although almost surely the metric space $\left(\mathbb{R}^d,T\right)$ is homeomorphic to the usual Euclidean $\mathbb{R}^d$, we prove that its Hausdorff dimension is given by $(γ-1)d/(γ-d)>d$, where $γ>d$ is a parameter of the model; which confirms a conjecture of Kahn. We also find that the metric space $\left(\mathbb{R}^d,T\right)$ equipped with the Lebesgue measure exhibits a multifractal property, as some points have untypically big balls around them.

math.PR

LSST: from Science Drivers to Reference Design and Anticipated Data Products

(Abridged) We describe here the most ambitious survey currently planned in the optical, the Large Synoptic Survey Telescope (LSST). A vast array of science will be enabled by a single wide-deep-fast sky survey, and LSST will have unique survey capability in the faint time domain. The LSST design is driven by four main science themes: probing dark energy and dark matter, taking an inventory of the Solar System, exploring the transient optical sky, and mapping the Milky Way. LSST will be a wide-field ground-based system sited at Cerro Pachón in northern Chile. The telescope will have an 8.4 m (6.5 m effective) primary mirror, a 9.6 deg$^2$ field of view, and a 3.2 Gigapixel camera. The standard observing sequence will consist of pairs of 15-second exposures in a given field, with two such visits in each pointing in a given night. With these repeats, the LSST system is capable of imaging about 10,000 square degrees of sky in a single filter in three nights. The typical 5$σ$ point-source depth in a single visit in $r$ will be $\sim 24.5$ (AB). The project is in the construction phase and will begin regular survey operations by 2022. The survey area will be contained within 30,000 deg$^2$ with $δ<+34.5^\circ$, and will be imaged multiple times in six bands, $ugrizy$, covering the wavelength range 320--1050 nm. About 90\% of the observing time will be devoted to a deep-wide-fast survey mode which will uniformly observe a 18,000 deg$^2$ region about 800 times (summed over all six bands) during the anticipated 10 years of operations, and yield a coadded map to $r\sim27.5$. The remaining 10\% of the observing time will be allocated to projects such as a Very Deep and Fast time domain survey. The goal is to make LSST data products, including a relational database of about 32 trillion observations of 40 billion objects, available to the public and scientists around the world.

astro-ph

The fast declining Type Ia supernova 2003gs, and evidence for a significant dispersion in near-infrared absolute magnitudes of fast decliners at maximum light

We obtained optical photometry of SN 2003gs on 49 nights, from 2 to 494 days after T(B_max). We also obtained near-IR photometry on 21 nights. SN 2003gs was the first fast declining Type Ia SN that has been well observed since SN 1999by. While it was subluminous in optical bands compared to more slowly declining Type Ia SNe, it was not subluminous at maximum light in the near-IR bands. There appears to be a bimodal distribution in the near-IR absolute magnitudes of Type Ia SNe at maximum light. Those that peak in the near-IR after T(B_max) are subluminous in the all bands. Those that peak in the near-IR prior to T(B_max), such as SN 2003gs, have effectively the same near-IR absolute magnitudes at maximum light regardless of the decline rate Delta m_15(B). Near-IR spectral evidence suggests that opacities in the outer layers of SN 2003gs are reduced much earlier than for normal Type Ia SNe. That may allow gamma rays that power the luminosity to escape more rapidly and accelerate the decline rate. This conclusion is consistent with the photometric behavior of SN 2003gs in the IR, which indicates a faster than normal decline from approximately normal peak brightness.

astro-ph.CO

Supernova progenitors and iron density evolution from SN rate evolution measurements

Using an extensive compilation of literature supernova rate data we study to which extent its evolution constrains the star formation history, the distribution of the type Ia supernova (SNIa) progenitor's lifetime, the mass range of core-collapse supernova (CCSN) progenitors, and the evolution of the iron density in the field. We find that the diagnostic power of the cosmic SNIa rate on their progenitor model is relatively weak. More promising is the use of the evolution of the SNIa rate in galaxy clusters. We find that the CCSN rate is compatible with a Salpeter IMF, with a minimum mass for their progenitors > 10 Msun. We estimate the evolution in the field of the iron density released by SNe and find that in the local universe the iron abundance should be ~ 0.1 solar. We discuss the difference between this value and the iron abundance in clusters.

astro-ph