SearcharxivSearch

arXiv subjects

Mathieu Galtier

Publications and source records attributed to Mathieu Galtier.

11 recordsLinked to original sources

Industry-Scale Orchestrated Federated Learning for Drug Discovery

To apply federated learning to drug discovery we developed a novel platform in the context of European Innovative Medicines Initiative (IMI) project MELLODDY (grant n{\deg}831472), which was comprised of 10 pharmaceutical companies, academic research labs, large industrial companies and startups. The MELLODDY platform was the first industry-scale platform to enable the creation of a global federated model for drug discovery without sharing the confidential data sets of the individual partners. The federated model was trained on the platform by aggregating the gradients of all contributing partners in a cryptographic, secure way following each training iteration. The platform was deployed on an Amazon Web Services (AWS) multi-account architecture running Kubernetes clusters in private subnets. Organisationally, the roles of the different partners were codified as different rights and permissions on the platform and administrated in a decentralized way. The MELLODDY platform generated new scientific discoveries which are described in a companion paper.

cs.LG

Collaborative Drug Discovery: Inference-level Data Protection Perspective

Pharmaceutical industry can better leverage its data assets to virtualize drug discovery through a collaborative machine learning platform. On the other hand, there are non-negligible risks stemming from the unintended leakage of participants' training data, hence, it is essential for such a platform to be secure and privacy-preserving. This paper describes a privacy risk assessment for collaborative modeling in the preclinical phase of drug discovery to accelerate the selection of promising drug candidates. After a short taxonomy of state-of-the-art inference attacks we adopt and customize several to the underlying scenario. Finally we describe and experiments with a handful of relevant privacy protection techniques to mitigate such attacks.

cs.CR

A deep learning architecture for temporal sleep stage classification using multivariate and multimodal time series

Sleep stage classification constitutes an important preliminary exam in the diagnosis of sleep disorders. It is traditionally performed by a sleep expert who assigns to each 30s of signal a sleep stage, based on the visual inspection of signals such as electroencephalograms (EEG), electrooculograms (EOG), electrocardiograms (ECG) and electromyograms (EMG). We introduce here the first deep learning approach for sleep stage classification that learns end-to-end without computing spectrograms or extracting hand-crafted features, that exploits all multivariate and multimodal Polysomnography (PSG) signals (EEG, EMG and EOG), and that can exploit the temporal context of each 30s window of data. For each modality the first layer learns linear spatial filters that exploit the array of sensors to increase the signal-to-noise ratio, and the last layer feeds the learnt representation to a softmax classifier. Our model is compared to alternative automatic approaches based on convolutional networks or decisions trees. Results obtained on 61 publicly available PSG records with up to 20 EEG channels demonstrate that our network architecture yields state-of-the-art performance. Our study reveals a number of insights on the spatio-temporal distribution of the signal of interest: a good trade-off for optimal classification performance measured with balanced accuracy is to use 6 EEG with 2 EOG (left and right) and 3 EMG chin channels. Also exploiting one minute of data before and after each data segment offers the strongest improvement when a limited number of channels is available. As sleep experts, our system exploits the multivariate and multimodal nature of PSG signals in order to deliver state-of-the-art classification performance with a small computational cost.

stat.ML

Morpheo: Traceable Machine Learning on Hidden data

Morpheo is a transparent and secure machine learning platform collecting and analysing large datasets. It aims at building state-of-the art prediction models in various fields where data are sensitive. Indeed, it offers strong privacy of data and algorithm, by preventing anyone to read the data, apart from the owner and the chosen algorithms. Computations in Morpheo are orchestrated by a blockchain infrastructure, thus offering total traceability of operations. Morpheo aims at building an attractive economic ecosystem around data prediction by channelling crypto-money from prediction requests to useful data and algorithms providers. Morpheo is designed to handle multiple data sources in a transfer learning approach in order to mutualize knowledge acquired from large datasets for applications with smaller but similar datasets.

cs.AI

Addressing nonlinearities in Monte Carlo

Monte Carlo is famous for accepting model extensions and model refinements up to infinite dimension. However, this powerful incremental design is based on a premise which has severely limited its application so far: a state-variable can only be recursively defined as a function of underlying state-variables if this function is linear. Here we show that this premise can be alleviated by projecting nonlinearities onto a polynomial basis and increasing the configuration space dimension. Considering phytoplankton growth in light-limited environments, radiative transfer in planetary atmospheres, electromagnetic scattering by particles, and concentrated solar power plant production, we prove the real-world usability of this advance in four test cases which were previously regarded as impracticable using Monte Carlo approaches. We also illustrate an outstanding feature of our method when applied to acute problems with interacting particles: handling rare events is now straightforward. Overall, our extension preserves the features that made the method popular: addressing nonlinearities does not compromise on model refinement or system complexity, and convergence rates remain independent of dimension. Published: Dauchet J, Bezian J-J, Blanco S, Caliot C, Charon J, Coustet C, El Hafi M, Eymet V, Farges O, Forest V, Fournier R, Galtier M, Gautrais J, Khuong A, Pelissier L, Piaud B, Roger M, Terr\'ee G, Weitz S (2018) Addressing nonlinearities in Monte Carlo. Sci. Rep. 8: 13302, DOI:10.1038/s41598-018-31574-4

physics.comp-ph

Ideomotor feedback control in a recurrent neural network

The architecture of a neural network controlling an unknown environment is presented. It is based on a randomly connected recurrent neural network from which both perception and action are simultaneously read and fed back. There are two concurrent learning rules implementing a sort of ideomotor control: (i) perception is learned along the principle that the network should predict reliably its incoming stimuli; (ii) action is learned along the principle that the prediction of the network should match a target time series. The coherent behavior of the neural network in its environment is a consequence of the interaction between the two principles. Numerical simulations show the promising performance of the approach, which can be turned into a local, and thus "biologically plausible", algorithm.

nlin.AO

Regular graphs maximize the variability of random neural networks

In this work we study the dynamics of systems composed of numerous interacting elements interconnected through a random weighted directed graph, such as models of random neural networks. We develop an original theoretical approach based on a combination of a classical mean-field theory originally developed in the context of dynamical spin-glass models, and the heterogeneous mean-field theory developed to study epidemic propagation on graphs. Our main result is that, surprisingly, increasing the variance of the in-degree distribution does not result in a more variable dynamical behavior, but on the contrary that the most variable behaviors are obtained in the regular graph setting. We further study how the dynamical complexity of the attractors is influenced by the statistical properties of the in-degree distribution.

nlin.CD

A local Echo State Property through the largest Lyapunov exponent

Echo State Networks are efficient time-series predictors, which highly depend on the value of the spectral radius of the reservoir connectivity matrix. Based on recent results on the mean field theory of driven random recurrent neural networks, enabling the computation of the largest Lyapunov exponent of an ESN, we develop a cheap algorithm to establish a local and operational version of the Echo State Property.

nlin.CD

Macroscopic equations governing noisy spiking neuronal populations

At functional scales, cortical behavior results from the complex interplay of a large number of excitable cells operating in noisy environments. Such systems resist to mathematical analysis, and computational neurosciences have largely relied on heuristic partial (and partially justified) macroscopic models, which successfully reproduced a number of relevant phenomena. The relationship between these macroscopic models and the spiking noisy dynamics of the underlying cells has since then been a great endeavor. Based on recent mean-field reductions for such spiking neurons, we present here {a principled reduction of large biologically plausible neuronal networks to firing-rate models, providing a rigorous} relationship between the macroscopic activity of populations of spiking neurons and popular macroscopic models, under a few assumptions (mainly linearity of the synapses). {The reduced model we derive consists of simple, low-dimensional ordinary differential equations with} parameters and {nonlinearities derived from} the underlying properties of the cells, and in particular the noise level. {These simple reduced models are shown to reproduce accurately the dynamics of large networks in numerical simulations}. Appropriate parameters and functions are made available {online} for different models of neurons: McKean, Fitzhugh-Nagumo and Hodgkin-Huxley models.

q-bio.NC

A biological gradient descent for prediction through a combination of STDP and homeostatic plasticity

Identifying, formalizing and combining biological mechanisms which implement known brain functions, such as prediction, is a main aspect of current research in theoretical neuroscience. In this letter, the mechanisms of Spike Timing Dependent Plasticity (STDP) and homeostatic plasticity, combined in an original mathematical formalism, are shown to shape recurrent neural networks into predictors. Following a rigorous mathematical treatment, we prove that they implement the online gradient descent of a distance between the network activity and its stimuli. The convergence to an equilibrium, where the network can spontaneously reproduce or predict its stimuli, does not suffer from bifurcation issues usually encountered in learning in recurrent neural networks.

q-bio.NC