SearcharxivSearch

arXiv subjects

Irene Balelli

Publications and source records attributed to Irene Balelli.

7 recordsLinked to original sources

Fed-BioMed: Open, Transparent and Trusted Federated Learning for Real-world Healthcare Applications

The real-world implementation of federated learning is complex and requires research and development actions at the crossroad between different domains ranging from data science, to software programming, networking, and security. While today several FL libraries are proposed to data scientists and users, most of these frameworks are not designed to find seamless application in medical use-cases, due to the specific challenges and requirements of working with medical data and hospital infrastructures. Moreover, governance, design principles, and security assumptions of these frameworks are generally not clearly illustrated, thus preventing the adoption in sensitive applications. Motivated by the current technological landscape of FL in healthcare, in this document we present Fed-BioMed: a research and development initiative aiming at translating federated learning (FL) into real-world medical research applications. We describe our design space, targeted users, domain constraints, and how these factors affect our current and future software architecture.

cs.LG

Fed-MIWAE: Federated Imputation of Incomplete Data via Deep Generative Models

Federated learning allows for the training of machine learning models on multiple decentralized local datasets without requiring explicit data exchange. However, data pre-processing, including strategies for handling missing data, remains a major bottleneck in real-world federated learning deployment, and is typically performed locally. This approach may be biased, since the subpopulations locally observed at each center may not be representative of the overall one. To address this issue, this paper first proposes a more consistent approach to data standardization through a federated model. Additionally, we propose Fed-MIWAE, a federated version of the state-of-the-art imputation method MIWAE, a deep latent variable model for missing data imputation based on variational autoencoders. MIWAE has the great advantage of being easily trainable with classical federated aggregators. Furthermore, it is able to deal with MAR (Missing At Random) data, a more challenging missing-data mechanism than MCAR (Missing Completely At Random), where the missingness of a variable can depend on the observed ones. We evaluate our method on multi-modal medical imaging data and clinical scores from a simulated federated scenario with the ADNI dataset. We compare Fed-MIWAE with respect to classical imputation methods, either performed locally or in a centralized fashion. Fed-MIWAE allows to achieve imputation accuracy comparable with the best centralized method, even when local data distributions are highly heterogeneous. In addition, thanks to the variational nature of Fed-MIWAE, our method is designed to perform multiple imputation, allowing for the quantification of the imputation uncertainty in the federated scenario.

stat.ML

A Differentially Private Probabilistic Framework for Modeling the Variability Across Federated Datasets of Heterogeneous Multi-View Observations

We propose a novel federated learning paradigm to model data variability among heterogeneous clients in multi-centric studies. Our method is expressed through a hierarchical Bayesian latent variable model, where client-specific parameters are assumed to be realization from a global distribution at the master level, which is in turn estimated to account for data bias and variability across clients. We show that our framework can be effectively optimized through expectation maximization (EM) over latent master's distribution and clients' parameters. We also introduce formal differential privacy (DP) guarantees compatibly with our EM optimization scheme. We tested our method on the analysis of multi-modal medical imaging data and clinical scores from distributed clinical datasets of patients affected by Alzheimer's disease. We demonstrate that our method is robust when data is distributed either in iid and non-iid manners, even when local parameters perturbation is included to provide DP guarantees. Moreover, the variability of data, views and centers can be quantified in an interpretable manner, while guaranteeing high-quality data reconstruction as compared to state-of-the-art autoencoding models and federated learning schemes. The code is available at https://gitlab.inria.fr/epione/federated-multi-views-ppca.

cs.LG

Parameter estimation in nonlinear mixed effect models based on ordinary differential equations: an optimal control approach

We present a parameter estimation method for nonlinear mixed effect models based on ordinary differential equations (NLME-ODEs). The method presented here aims at regularizing the estimation problem in presence of model misspecifications, practical identifiability issues and unknown initial conditions. For doing so, we define our estimator as the minimizer of a cost function which incorporates a possible gap between the assumed model at the population level and the specific individual dynamic. The cost function computation leads to formulate and solve optimal control problems at the subject level. This control theory approach allows to bypass the need to know or estimate initial conditions for each subject and it regularizes the estimation problem in presence of poorly identifiable parameters. Comparing to maximum likelihood, we show on simulation examples that our method improves estimation accuracy in possibly partially observed systems with unknown initial conditions or poorly identifiable parameters with or without model error. We conclude this work with a real application on antibody concentration data after vaccination against Ebola virus coming from phase 1 trials. We use the estimated model discrepancy at the subject level to analyze the presence of model misspecification.

stat.ME

Multi-type Galton-Watson processes with affinity-dependent selection applied to antibody affinity maturation

We analyze the interactions between division, mutation and selection in a simplified evolutionary model, assuming that the population observed can be classified into fitness levels. The construction of our mathematical framework is motivated by the modeling of antibody affinity maturation of B-cells in Germinal Centers during an immune response. This is a key process in adaptive immunity leading to the production of high affinity antibodies against a presented antigen. Our aim is to understand how the different biological parameters affect the system's functionality. We identify the existence of an optimal value of the selection rate, able to maximize the number of selected B-cells for a given generation.

math.PR

Branching Random Walks on Binary Strings for Evolutionary Processes

In this article, we study branching random walks on graphs modeling division-mutation processes inspired by adaptive immunity. We apply the theory of expander graphs on mutation rules in evolutionary processes and obtain estimates for the cover times of the branching random walks. This analysis reveals an unexpected saturation phenomenon : increasing the mutation rate above a certain threshold does not enhance the speed of state-space exploration.

math.PR

Random walks on binary strings applied to the somatic hypermutation of B-cells

Within the germinal center in follicles, B-cells proliferate, mutate and differentiate, while being submitted to a powerful selection~: a micro-evolutionary mechanism at the heart of adaptive immunity. A new foreign pathogen is confronted to our immune system, the mutation mechanism that allows B-cells to adapt to it is called {\em somatic hypermutation}~: a programmed process of mutation affecting B-cell receptors at extremely high rate. By considering random walks on graphs, we introduce and analyze a simplified mathematical model in order to understand this extremely efficient learning process. The structure of the graph reflects the choice of the mutation rule. We focus on the impact of this choice on typical time-scales of the graphs' exploration. We derive explicit formulas to evaluate the expected hitting time to cover a given Hamming distance on the graphs under consideration. This characterizes the efficiency of these processes in driving antibody affinity maturation. In a further step we present a biologically more involved model and discuss its numerical outputs within our mathematical framework. We provide as well limitations and possible extensions of our approach.

math.PR