SearcharxivSearch

arXiv subjects

Jean-Philippe Thiran

Publications and source records attributed to Jean-Philippe Thiran.

At least 19 recordsLinked to original sources

FeatEHR-LLM: Leveraging Large Language Models for Feature Engineering in Electronic Health Records

Feature engineering for Electronic Health Records (EHR) is complicated by irregular observation intervals, variable measurement frequencies, and structural sparsity inherent to clinical time series. Existing automated methods either lack clinical domain awareness or assume clean, regularly sampled inputs, limiting their applicability to real-world EHR data. We present \textbf{FeatEHR-LLM}, a framework that leverages Large Language Models (LLMs) to generate clinically meaningful tabular features from irregularly sampled EHR time series. To limit patient privacy exposure, the LLM operates exclusively on dataset schemas and task descriptions rather than raw patient records. A tool-augmented generation mechanism equips the LLM with specialized routines for querying irregular temporal data, enabling it to produce executable feature-extraction code that explicitly handles uneven observation patterns and informative sparsity. FeatEHR-LLM supports both univariate and multivariate feature generation through an iterative, validation-in-the-loop pipeline. Evaluated on eight clinical prediction tasks across four ICU datasets, our framework achieves the highest mean AUROC on 7 out of 8 tasks, with improvements of up to 6 percentage points over strong baselines. Code is available at github.com/hojjatkarami/FeatEHR-LLM.

cs.LG

Exact analytical PGSE signal for diffusion confined to a cylindrical surface using a spectral Laplacian formalism

Pulsed-gradient spin-echo (PGSE) MRI experiments probe molecular self-diffusion through spin phase accumulation under time-dependent magnetic field gradients. For diffusion confined to cylindrical surfaces, existing analytical signal models typically rely on the narrow-pulse limit, approximate treatments of finite gradient durations, or the Gaussian phase approximation, which become increasingly inaccurate at high diffusion weightings. Here, we derive an exact analytical solution of the Bloch-Torrey equation for diffusion confined to a cylindrical surface under finite PGSE gradients and obtain the corresponding diffusion MRI signal expression valid for arbitrary gradient durations and separations. The derivation is based on a spectral matrix formalism of the Laplace operator in the eigenbasis of the confining geometry. The signal is expressed as a product of non-commuting matrix exponentials, without approximations to the diffusion propagator or the spin phase distribution. We further introduce a reduced real spectral basis exploiting the symmetry of the cylindrical surface, substantially improving computational efficiency. Building on this exact formulation, we develop efficient numerical strategies for repeated signal evaluations, including a Strang splitting approximation of the matrix exponentials and an efficient computation of the spherical mean signal using Gauss-Legendre quadrature. The analytical signal is validated against Monte Carlo simulations over a wide range of cylinder radii and experimental parameters. The accelerated implementations are benchmarked against the exact formulation to quantify accuracy-runtime trade-offs. These results establish a computationally efficient framework for evaluating directional and orientationally averaged diffusion MRI signals in applications requiring large numbers of model evaluations.

physics.med-ph

Spatially Regularized Super-Resolved Constrained Spherical Deconvolution (SR$^2$-CSD) of Diffusion MRI Data

Constrained Spherical Deconvolution (CSD) is widely used to estimate the white matter fiber orientation distribution (FOD) from diffusion MRI data. Its angular resolution depends on the maximum spherical harmonic order ($l_{max}$): low $l_{max}$ yields smooth but poorly resolved FODs, while high $l_{max}$, as in Super-CSD, enables resolving fiber crossings with small inter-fiber angles but increases sensitivity to noise. In this proof-of-concept study, we introduce Spatially Regularized Super-Resolved CSD (SR$^2$-CSD), a novel method that regularizes Super-CSD using a spatial FOD prior estimated via a self-calibrated total variation denoiser. We evaluated SR$^2$-CSD against CSD and Super-CSD across four datasets: (i) the HARDI-2013 challenge numerical phantom, assessing angular and peak number errors across multiple signal-to-noise ratio (SNR) levels and CSD variants (single-/multi-shell, single-/multi-tissue); (ii) the Sherbrooke in vivo dataset, evaluating spatial coherence of FODs; (iii) a six-subject test-retest dataset acquired with both full (96 gradient directions) and subsampled (45 directions) protocols, assessing reproducibility; and (iv) the DiSCo phantom, evaluating tractography accuracy under varying SNR levels and multiple noise repetitions. Across all evaluations, SR$^2$-CSD consistently reduced angular and peak number errors, improved spatial coherence, enhanced test-retest reproducibility, and yielded connectivity matrices more strongly correlated with ground-truth. Most improvements were statistically significant under multiple-comparison correction. These results demonstrate that incorporating spatial priors into CSD is feasible, mitigates estimation instability, and improves FOD reconstruction accuracy.

physics.med-ph

A Simple Framework for Open-Vocabulary Zero-Shot Segmentation

Zero-shot classification capabilities naturally arise in models trained within a vision-language contrastive framework. Despite their classification prowess, these models struggle in dense tasks like zero-shot open-vocabulary segmentation. This deficiency is often attributed to the absence of localization cues in captions and the intertwined nature of the learning process, which encompasses both image representation learning and cross-modality alignment. To tackle these issues, we propose SimZSS, a Simple framework for open-vocabulary Zero-Shot Segmentation. The method is founded on two key principles: i) leveraging frozen vision-only models that exhibit spatial awareness while exclusively aligning the text encoder and ii) exploiting the discrete nature of text and linguistic knowledge to pinpoint local concepts within captions. By capitalizing on the quality of the visual representations, our method requires only image-caption pairs datasets and adapts to both small curated and large-scale noisy datasets. When trained on COCO Captions across 8 GPUs, SimZSS achieves state-of-the-art results on 7 out of 8 benchmark datasets in less than 15 minutes.

cs.CV

ReservoirTTA: Prolonged Test-time Adaptation for Evolving and Recurring Domains

This paper introduces ReservoirTTA, a novel plug-in framework designed for prolonged test-time adaptation (TTA) in scenarios where the test domain continuously shifts over time, including cases where domains recur or evolve gradually. At its core, ReservoirTTA maintains a reservoir of domain-specialized models -- an adaptive test-time model ensemble -- that both detects new domains via online clustering over style features of incoming samples and routes each sample to the appropriate specialized model, and thereby enables domain-specific adaptation. This multi-model strategy overcomes key limitations of single model adaptation, such as catastrophic forgetting, inter-domain interference, and error accumulation, ensuring robust and stable performance on sustained non-stationary test distributions. Our theoretical analysis reveals key components that bound parameter variance and prevent model collapse, while our plug-in TTA module mitigates catastrophic forgetting of previously encountered domains. Extensive experiments on scene-level corruption benchmarks (ImageNet-C, CIFAR-10/100-C), object-level style shifts (DomainNet-126, PACS), and semantic segmentation (Cityscapes->ACDC) covering recurring and continuously evolving domain shifts -- show that ReservoirTTA substantially improves adaptation accuracy and maintains stable performance across prolonged, recurring shifts, outperforming state-of-the-art methods. Our code is publicly available at https://github.com/LTS5/ReservoirTTA.

cs.CV

Towards Early Detection: AI-Based Five-Year Forecasting of Breast Cancer Risk Using Digital Breast Tomosynthesis Imaging

As early detection of breast cancer strongly favors successful therapeutic outcomes, there is major commercial interest in optimizing breast cancer screening. However, current risk prediction models achieve modest performance and do not incorporate digital breast tomosynthesis (DBT) imaging, which was FDA-approved for breast cancer screening in 2011. To address this unmet need, we present a deep learning (DL)-based framework capable of forecasting an individual patient's 5-year breast cancer risk directly from screening DBT. Using an unparalleled dataset of 161,753 DBT examinations from 50,590 patients, we trained a risk predictor based on features extracted using the Meta AI DINOv2 image encoder, combined with a cumulative hazard layer, to assess a patient's likelihood of developing breast cancer over five years. On a held-out test set, our best-performing model achieved an AUROC of 0.80 on predictions within 5 years. These findings reveal the high potential of DBT-based DL approaches to complement traditional risk assessment tools, and serve as a promising basis for additional investigation to validate and enhance our work.

eess.IV

Uncertainty modeling for fine-tuned implicit functions

Implicit functions such as Neural Radiance Fields (NeRFs), occupancy networks, and signed distance functions (SDFs) have become pivotal in computer vision for reconstructing detailed object shapes from sparse views. Achieving optimal performance with these models can be challenging due to the extreme sparsity of inputs and distribution shifts induced by data corruptions. To this end, large, noise-free synthetic datasets can serve as shape priors to help models fill in gaps, but the resulting reconstructions must be approached with caution. Uncertainty estimation is crucial for assessing the quality of these reconstructions, particularly in identifying areas where the model is uncertain about the parts it has inferred from the prior. In this paper, we introduce Dropsembles, a novel method for uncertainty estimation in tuned implicit functions. We demonstrate the efficacy of our approach through a series of experiments, starting with toy examples and progressing to a real-world scenario. Specifically, we train a Convolutional Occupancy Network on synthetic anatomical data and test it on low-resolution MRI segmentations of the lumbar spine. Our results show that Dropsembles achieve the accuracy and calibration levels of deep ensembles but with significantly less computational cost.

cs.CV

Slide-Level Prompt Learning with Vision Language Models for Few-Shot Multiple Instance Learning in Histopathology

In this paper, we address the challenge of few-shot classification in histopathology whole slide images (WSIs) by utilizing foundational vision-language models (VLMs) and slide-level prompt learning. Given the gigapixel scale of WSIs, conventional multiple instance learning (MIL) methods rely on aggregation functions to derive slide-level (bag-level) predictions from patch representations, which require extensive bag-level labels for training. In contrast, VLM-based approaches excel at aligning visual embeddings of patches with candidate class text prompts but lack essential pathological prior knowledge. Our method distinguishes itself by utilizing pathological prior knowledge from language models to identify crucial local tissue types (patches) for WSI classification, integrating this within a VLM-based MIL framework. Our approach effectively aligns patch images with tissue types, and we fine-tune our model via prompt learning using only a few labeled WSIs per category. Experimentation on real-world pathological WSI datasets and ablation studies highlight our method's superior performance over existing MIL- and VLM-based methods in few-shot WSI classification tasks. Our code is publicly available at https://github.com/LTS5/SLIP.

cs.CV

What to align in multimodal contrastive learning?

Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive MultiModal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on the seven multimodal benchmarks. Code is available at https://github.com/Duplums/CoMM

cs.LG

A diffusion MRI model for random walks confined on cylindrical surfaces: Towards non-invasive quantification of myelin sheath radius

Quantifying the myelin sheath radius of myelinated axons in vivo is important for understanding, diagnosing, and monitoring various neurological disorders. Despite advancements in diffusion MRI (dMRI) microstructure techniques, there are currently no models specifically designed to estimate myelin sheath radii. This proof-of-concept theoretical study presents two novel dMRI models that characterize the signal from water diffusion confined to cylindrical surfaces, approximating myelin water diffusion. We derive their spherical mean signals, eliminating fiber orientation and dispersion effects for convenience. These models are further extended to account for multiple concentric cylinders, mimicking the layered structure of myelin. Additionally, we introduce a method to convert histological distributions of axonal inner radii from the literature into myelin sheath radius distributions. We also derive analytical expressions to estimate the effective myelin sheath radius expected from these distributions. Monte Carlo (MC) simulations conducted in cylindrical and spiral geometries validate the models. These simulations demonstrate agreement with analytical predictions. Furthermore, we observe significant correlations between the effective radii derived from histological distributions and those obtained by fitting the dMRI signal to a single-cylinder model. These models may be integrated with existing multi-compartment dMRI techniques, opening the door to non-invasive in vivo assessments of myelin sheath radii. Such assessments would require MRI scanners equipped with strong diffusion gradients, allowing measurements with short echo times. Further work is required to validate the technique with real dMRI data and histological measurements.

physics.med-ph

Ground-truth effects in learning-based fiber orientation distribution estimation in neonatal brains

Diffusion Magnetic Resonance Imaging (dMRI) is a non-invasive method for depicting brain microstructure in vivo. Fiber orientation distributions (FODs) are mathematical representations extensively used to map white matter fiber configurations. Recently, FOD estimation with deep neural networks has seen growing success, in particular, those of neonates estimated with fewer diffusion measurements. These methods are mostly trained on target FODs reconstructed with multi-shell multi-tissue constrained spherical deconvolution (MSMT-CSD), which might not be the ideal ground truth for developing brains. Here, we investigate this hypothesis by training a state-of-the-art model based on the U-Net architecture on both MSMT-CSD and single-shell three-tissue constrained spherical deconvolution (SS3T-CSD). Our results suggest that SS3T-CSD might be more suited for neonatal brains, given that the ratio between single and multiple fiber-estimated voxels with SS3T-CSD is more realistic compared to MSMT-CSD. Additionally, increasing the number of input gradient directions significantly improves performance with SS3T-CSD over MSMT-CSD. Finally, in an age domain-shift setting, SS3T-CSD maintains robust performance across age groups, indicating its potential for more accurate neonatal brain imaging.

eess.IV

Cross-Age and Cross-Site Domain Shift Impacts on Deep Learning-Based White Matter Fiber Estimation in Newborn and Baby Brains

Deep learning models have shown great promise in estimating tissue microstructure from limited diffusion magnetic resonance imaging data. However, these models face domain shift challenges when test and train data are from different scanners and protocols, or when the models are applied to data with inherent variations such as the developing brains of infants and children scanned at various ages. Several techniques have been proposed to address some of these challenges, such as data harmonization or domain adaptation in the adult brain. However, those techniques remain unexplored for the estimation of fiber orientation distribution functions in the rapidly developing brains of infants. In this work, we extensively investigate the age effect and domain shift within and across two different cohorts of 201 newborns and 165 babies using the Method of Moments and fine-tuning strategies. Our results show that reduced variations in the microstructural development of babies in comparison to newborns directly impact the deep learning models' cross-age performance. We also demonstrate that a small number of target domain samples can significantly mitigate domain shift problems.

eess.IV

PIV3CAMS: a multi-camera dataset for multiple computer vision problems and its application to novel view-point synthesis

The modern approaches for computer vision tasks significantly rely on machine learning, which requires a large number of quality images. While there is a plethora of image datasets with a single type of images, there is a lack of datasets collected from multiple cameras. In this thesis, we introduce Paired Image and Video data from three CAMeraS, namely PIV3CAMS, aimed at multiple computer vision tasks. The PIV3CAMS dataset consists of 8385 pairs of images and 82 pairs of videos taken from three different cameras: Canon D5 Mark IV, Huawei P20, and ZED stereo camera. The dataset includes various indoor and outdoor scenes from different locations in Zurich (Switzerland) and Cheonan (South Korea). Some of the computer vision applications that can benefit from the PIV3CAMS dataset are image/video enhancement, view interpolation, image matching, and much more. We provide a careful explanation of the data collection process and detailed analysis of the data. The second part of this thesis studies the usage of depth information in the view synthesizing task. In addition to the regeneration of a current state-of-the-art algorithm, we investigate several proposed alternative models that integrate depth information geometrically. Through extensive experiments, we show that the effect of depth is crucial in small view changes. Finally, we apply our model to the introduced PIV3CAMS dataset to synthesize novel target views as an example application of PIV3CAMS.

cs.CV

Un-Mixing Test-Time Normalization Statistics: Combatting Label Temporal Correlation

Recent test-time adaptation methods heavily rely on nuanced adjustments of batch normalization (BN) parameters. However, one critical assumption often goes overlooked: that of independently and identically distributed (i.i.d.) test batches with respect to unknown labels. This oversight leads to skewed BN statistics and undermines the reliability of the model under non-i.i.d. scenarios. To tackle this challenge, this paper presents a novel method termed 'Un-Mixing Test-Time Normalization Statistics' (UnMix-TNS). Our method re-calibrates the statistics for each instance within a test batch by mixing it with multiple distinct statistics components, thus inherently simulating the i.i.d. scenario. The core of this method hinges on a distinctive online unmixing procedure that continuously updates these statistics components by incorporating the most similar instances from new test batches. Remarkably generic in its design, UnMix-TNS seamlessly integrates with a wide range of leading test-time adaptation methods and pre-trained architectures equipped with BN layers. Empirical evaluations corroborate the robustness of UnMix-TNS under varied scenarios-ranging from single to continual and mixed domain shifts, particularly excelling with temporally correlated test data and corrupted non-i.i.d. real-world streams. This adaptability is maintained even with very small batch sizes or single instances. Our results highlight UnMix-TNS's capacity to markedly enhance stability and performance across various benchmarks. Our code is publicly available at https://github.com/devavratTomar/unmixtns.

cs.CV

CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping

Leveraging nearest neighbor retrieval for self-supervised representation learning has proven beneficial with object-centric images. However, this approach faces limitations when applied to scene-centric datasets, where multiple objects within an image are only implicitly captured in the global representation. Such global bootstrapping can lead to undesirable entanglement of object representations. Furthermore, even object-centric datasets stand to benefit from a finer-grained bootstrapping approach. In response to these challenges, we introduce a novel Cross-Image Object-Level Bootstrapping method tailored to enhance dense visual representation learning. By employing object-level nearest neighbor bootstrapping throughout the training, CrIBo emerges as a notably strong and adequate candidate for in-context learning, leveraging nearest neighbor retrieval at test time. CrIBo shows state-of-the-art performance on the latter task while being highly competitive in more standard downstream segmentation tasks. Our code and pretrained models are publicly available at https://github.com/tileb1/CrIBo.

cs.CV

Distill-SODA: Distilling Self-Supervised Vision Transformer for Source-Free Open-Set Domain Adaptation in Computational Pathology

Developing computational pathology models is essential for reducing manual tissue typing from whole slide images, transferring knowledge from the source domain to an unlabeled, shifted target domain, and identifying unseen categories. We propose a practical setting by addressing the above-mentioned challenges in one fell swoop, i.e., source-free open-set domain adaptation. Our methodology focuses on adapting a pre-trained source model to an unlabeled target dataset and encompasses both closed-set and open-set classes. Beyond addressing the semantic shift of unknown classes, our framework also deals with a covariate shift, which manifests as variations in color appearance between source and target tissue samples. Our method hinges on distilling knowledge from a self-supervised vision transformer (ViT), drawing guidance from either robustly pre-trained transformer models or histopathology datasets, including those from the target domain. In pursuit of this, we introduce a novel style-based adversarial data augmentation, serving as hard positives for self-training a ViT, resulting in highly contextualized embeddings. Following this, we cluster semantically akin target images, with the source model offering weak pseudo-labels, albeit with uncertain confidence. To enhance this process, we present the closed-set affinity score (CSAS), aiming to correct the confidence levels of these pseudo-labels and to calculate weighted class prototypes within the contextualized embedding space. Our approach establishes itself as state-of-the-art across three public histopathological datasets for colorectal cancer assessment. Notably, our self-training method seamlessly integrates with open-set detection methods, resulting in enhanced performance in both closed-set and open-set recognition tasks.

cs.CV

Pore size estimation in axon-mimicking microfibres with diffusion-relaxation MRI

Purpose: This study aims to evaluate two distinct approaches for fibre radius estimation using diffusion-relaxation MRI data acquired in biomimetic microfibre phantoms that mimic hollow axons. The methods considered are the spherical mean power-law approach and a T2-based pore size estimation technique. Theory and Methods: A general diffusion-relaxation theoretical model for the spherical mean signal from water molecules within a distribution of cylinders with varying radii was introduced, encompassing the evaluated models as particular cases. Additionally, a new numerical approach was presented for estimating effective radii (i.e., MRI-visible mean radii) from the ground truth radii distributions, not reliant on previous theoretical approximations and adaptable to various acquisition sequences. The ground truth radii were obtained from Scanning Electron Microscope images. Results: Both methods show a linear relationship between effective radii estimated from MRI data and ground-truth radii distributions, though some discrepancies were observed. The spherical mean power-law method overestimated fibre radii. Conversely, the T2-based method exhibited higher sensitivity to smaller fibre radii but faced limitations in accurately estimating the radius in one particular phantom, possibly due to material-specific relaxation changes. Conclusion: The study demonstrates the feasibility of both techniques to predict pore sizes of hollow microfibres. The T2-based technique, unlike the spherical mean power-law method, does not demand ultra-high diffusion gradients but requires calibration with known radius distributions. This research contributes to the ongoing development and evaluation of neuroimaging techniques for fibre radius estimation, highlights the advantages and limitations of both methods and provides datasets for reproducible research.

physics.med-ph

Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning

Most self-supervised methods for representation learning leverage a cross-view consistency objective i.e., they maximize the representation similarity of a given image's augmented views. Recent work NNCLR goes beyond the cross-view paradigm and uses positive pairs from different images obtained via nearest neighbor bootstrapping in a contrastive setting. We empirically show that as opposed to the contrastive learning setting which relies on negative samples, incorporating nearest neighbor bootstrapping in a self-distillation scheme can lead to a performance drop or even collapse. We scrutinize the reason for this unexpected behavior and provide a solution. We propose to adaptively bootstrap neighbors based on the estimated quality of the latent space. We report consistent improvements compared to the naive bootstrapping approach and the original baselines. Our approach leads to performance improvements for various self-distillation method/backbone combinations and standard downstream tasks. Our code is publicly available at https://github.com/tileb1/AdaSim.

cs.CV