SearcharxivSearch

arXiv subjects

Federico Felizzi

Publications and source records attributed to Federico Felizzi.

7 recordsLinked to original sources

Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification

Structured tabular data dominates clinical medicine, yet existing benchmarks fail to reflect real-world properties like complex survey sampling, demographic oversampling, and subgroup fairness. We introduce the NHANES Accelerometry Cardiometabolic Benchmark, derived from NHANES 2003-2006, comprising 1,381 adults with hip-worn accelerometry, fasting laboratory biomarkers, dietary intake, and anthropometrics. We evaluate three tabular learning methods -- ridge regression, XGBoost, and the foundation model TabPFN v2 -- to predict glycated haemoglobin (HbA1c), fasting triglycerides, and C-reactive protein (CRP) from activity phenotypes and lifestyle covariates. TabPFN v2 achieves the best overall performance (HbA1c R^2=0.156, CRP R^2=0.383), while triglycerides remain largely unpredictable (R^2 < 0.05), consistent with known genetic dominance. We apply split conformal prediction to generate distribution-free 90% prediction intervals and evaluate demographic coverage equity across sex and race/ethnicity subgroups. Marginal coverage aligns with the 90% target for CRP and HbA1c but falls below for triglycerides. At the subgroup level, we observe localized undercoverage (e.g., HbA1c for Mexican American participants), illustrating the gap between marginal guarantees and the conditional coverage required for clinical fairness. Code and data are at https://github.com/felizzi/nhanes-accel-cardiometabolic-benchmark.

cs.LG

EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation

While Large Language Models (LLMs) have demonstrated high proficiency on English-centric medical examinations, their performance often declines when faced with non-English languages and multimodal diagnostic tasks. This study protocol describes the development of EuropeMedQA, the first comprehensive, multilingual, and multimodal medical examination dataset sourced from official regulatory exams in Italy, France, Spain, and Portugal. Following FAIR data principles and SPIRIT-AI guidelines, we describe a rigorous curation process and an automated translation pipeline for comparative analysis. We evaluate contemporary multimodal LLMs using a zero-shot, strictly constrained prompting strategy to assess cross-lingual transfer and visual reasoning. EuropeMedQA aims to provide a contamination-resistant benchmark that reflects the complexity of European clinical practices and fosters the development of more generalizable medical AI.

cs.CL

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

Large vision language models (VLMs) have achieved impressive performance on medical visual question answering benchmarks, yet their reliance on visual information remains unclear. We investigate whether frontier VLMs demonstrate genuine visual grounding when answering Italian medical questions by testing four state-of-the-art models: Claude Sonnet 4.5, GPT-4o, GPT-5-mini, and Gemini 2.0 flash exp. Using 60 questions from the EuropeMedQA Italian dataset that explicitly require image interpretation, we substitute correct medical images with blank placeholders to test whether models truly integrate visual and textual information. Our results reveal striking variability in visual dependency: GPT-4o shows the strongest visual grounding with a 27.9pp accuracy drop (83.2% [74.6%, 91.7%] to 55.3% [44.1%, 66.6%]), while GPT-5-mini, Gemini, and Claude maintain high accuracy with modest drops of 8.5pp, 2.4pp, and 5.6pp respectively. Analysis of model-generated reasoning reveals confident explanations for fabricated visual interpretations across all models, suggesting varying degrees of reliance on textual shortcuts versus genuine visual analysis. These findings highlight critical differences in model robustness and the need for rigorous evaluation before clinical deployment.

cs.CV

Economic impact of biomarker-based aging interventions on healthcare costs and individual value

We investigate the economic impact of controlling the pace of aging through biomarker monitoring and targeted interventions. Using the DunedinPACE epigenetic clock as a measure of biological aging rate, we model how different intervention scenarios affect frailty trajectories and their subsequent influence on healthcare costs, lifespan, and health quality. Our model demonstrates that controlling DunedinPACE from age 50 onwards can reduce frailty prevalence, resulting in cumulative healthcare savings of up to CHF 131,608 per person over 40 years in our most optimistic scenario. From an individual perspective, the willingness to pay for such interventions reaches CHF 6.7 million when accounting for both extended lifespan and improved health quality. These findings suggest substantial economic value in technologies that can monitor and modify biological aging rates, providing evidence for both healthcare systems and consumer-focused business models in longevity medicine.

q-bio.QM

Spatial organization of proteomes: A low-rank approximation

We investigate the problem of signal transduction via a descriptive analysis of the spatial organization of the complement of proteins exerting a certain function within a cellular compartment. We propose a scheme to assign a numerical value to individual proteins in a protein interaction network by means of a simple optimization algorithm. We test our procedure against datasets focusing on the proteomes in the neurite and soma compartments.

q-bio.MN

Microtubule tracking from stochastic optical reconstruction microscopy images

Our work aims at using quantitative imaging tools to complement the limitation of noise encountered by high resolution fluorescence microscopy methods. Several cycles of fluorophore activation, imaging and deactivation produce a sequence of images in which the signals of individual fluorophores do not overlap, due to the low light intensity during their activation. The centroid position of each fluorophore is then determined by Gaussian fitting of each signal, where the final resolution depends on the precision with which each fluorphore is localized. Superimposing the images will result in having the same fluorophore mapped onto a `cloud' of locations. The most significant information of the superimposed images is contained in the macro-structures identifying microtubules, mitochondria or other organelles. Cascades of binary image processing algorithms are applied in order to isolate the larger organelles. A Markovian algorithm selecting the nearest neighbour is finally applied to the de-noised images, to automatically extract relevant information on microtubules. Our work supplements advancements in experimental technologies with computational methods, helping quantifying sub-cellular properties with high accuracy.

q-bio.QM

A Monte Carlo study of ligand-dependent integrin signal initiation

Integrins are allosteric cell adhesion receptors that control many important processes, including cell migration, proliferation, and apoptosis. Ligand binding activates integrins by stabilizing an integrin conformation with separated cytoplasmic tails, thus enabling the binding of proteins that mediate cytoplasmic signaling. Experiments demonstrate a high sensitivity of integrin signaling to ligand density and this has been accounted mainly to avidity effects. Based on experimental data we have developed a quantitative Monte Carlo model for integrin signal initiation. We show that within the physiological ligand density range avidity effects cannot explain the sensitivity of cellular signaling to small changes in ligand density. Src kinases are among the first proteins to be activated, possibly by trans auto-phosphorylation. We calculate the extent of integrin and ligand clustering as well as the speed and extent of Src kinase activation by trans auto-phosphorylation or direct binding at different experimentally monitored ligand densities. We find that the experimentally observed ligand density dependency can be reproduced if Src kinases are activated by trans auto-phosphorylation or some other mechanism limits integrin-dependent Src kinase activation. We propose that Src kinase and thus cell activation by trans auto-phosphorylation may provide a mechanism to enable ligand-density dependent responses at physiological ligand densities. The capacity to detect small differences in ligand density at a ligand density that is large enough to permit cell adhesion is likely to be important for haptotaxis.

q-bio.QM