SearcharxivSearch

arXiv subjects

Pierre Pudlo

Publications and source records attributed to Pierre Pudlo.

At least 19 recordsLinked to original sources

Owl-z: a Bayesian tool to select z \geq 7 quasars

This paper presents Owl-z, a Bayesian code aiming at identifying z \geq 7 quasars in wide field optical and near-infrared surveys. By construction,the code can also be used to select objects that contaminate the high-z quasar population, i.e. brown dwarfs and early-type galaxies at intermediate redshifts. The code can be adapted for the selection of high-z galaxies, and although it has been tuned to the Euclid Wide Survey, it can be easily adapted to other photometric surveys. The code input data are the object's photometric data and its galactic longitude and latitude, and the code output data are the probabilities of the modelled populations of high-z quasars, brown dwarfs and early-type galaxies at intermediate redshift. As part of the validation, Owl-z could re-identify all spectroscopically confirmed quasars at z \geq 7, demonstrating the code's versatility in applying to different photometric catalogues. The performance of Owl-z, based on a metric combining completeness and purity called F-measure, is analysed in the case of Euclid using simulated data in a wide range of redshifts (7 \leq z \leq 12) and H-band Euclid magnitudes (18 \leq H_E \leq 24.5). The results show that Owl-z reaches full performance for bright sources (H_E \lesssim 22), independently of the redshift. We show that the probability threshold used to select promising quasar candidates can be adjusted after processing to fine-tune the F-measure value of candidates depending on their magnitude and redshift estimates. We show that for objects brighter than about two magnitudes above the survey detection limit, Owl-z provides a classification that will facilitate the optimisation of photometric and spectroscopic confirmation campaigns. In conclusion, Owl-z is a powerful public tool to help select high-z quasars, brown dwarfs or early-type galaxies at intermediate redshifts in Euclid or other wide-field surveys.

astro-ph.CO

Reconciling Binary Replicates: Beyond the Average

Binary observations are often repeated to improve data quality, creating technical replicates. Several scoring methods are commonly used to infer the actual individual state and obtain a probability for each state. The common practice of averaging replicates has limitations, and alternative methods for scoring and classifying individuals are proposed. Additionally, an indecisive response might be wiser than classifying all individuals based on their replicates in the medical context, where 1 indicates a particular health condition. Building on the inherent limitations of the averaging approach, three alternative methods are examined: the median, maximum penalized likelihood estimation, and a Bayesian algorithm. The theoretical analysis suggests that the proposed alternatives outperform the averaging approach, especially the Bayesian method, which incorporates uncertainty and provides credible intervals. Simulations and real-world medical datasets are used to demonstrate the practical implications of these methods for improving diagnostic accuracy and disease prevalence estimation.

stat.ME

Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations

Based on audio recordings made once a month during the first 12 months of a child's life, we propose a new method for clustering this set of vocalizations. We use a topologically augmented representation of the vocalizations, employing two persistence diagrams for each vocalization: one computed on the surface of its spectrogram and one on the Takens' embeddings of the vocalization. A synthetic persistent variable is derived for each diagram and added to the MFCCs (Mel-frequency cepstral coefficients). Using this representation, we fit a non-parametric Bayesian mixture model with a Dirichlet process prior to model the number of components. This procedure leads to a novel data-driven categorization of vocal productions. Our findings reveal the presence of 8 clusters of vocalizations, allowing us to compare their temporal distribution and acoustic profiles in the first 12 months of life.

stat.AP

A continuous approach of modeling tumorigenesis and axons regulation for the pancreatic cancer

The pancreatic innervation undergoes dynamic remodeling during the development of pancreatic ductal adenocarcinoma (PDAC). Denervation experiments have shown that different types of axons can exert either pro- or anti-tumor effects, but conflicting results exist in the literature, leaving the overall influence of the nervous system on PDAC incompletely understood. To address this gap, we propose a continuous mathematical model of nerve-tumor interactions that allows in silico simulation of denervation at different phases of tumor development. This model takes into account the pro- or anti-tumor properties of different types of axons (sympathetic or sensory) and their distinct remodeling dynamics during PDAC development. We observe a "shift effect" where an initial pro-tumor effect of sympathetic axon denervation is later outweighed by the anti-tumor effect of sensory axon denervation, leading to a transition from an overall protective to a deleterious role of the nervous system on PDAC tumorigenesis. Our model also highlights the importance of the impact of sympathetic axon remodeling dynamics on tumor progression. These findings may guide strategies targeting the nervous system to improve PDAC treatment.

math.AP

Topological data analysis of human vowels: Persistent homologies across representation spaces

Topological Data Analysis (TDA) has been successfully used for various tasks in signal/image processing, from visualization to supervised/unsupervised classification. Often, topological characteristics are obtained from persistent homology theory. The standard TDA pipeline starts from the raw signal data or a representation of it. Then, it consists in building a multiscale topological structure on the top of the data using a pre-specified filtration, and finally to compute the topological signature to be further exploited. The commonly used topological signature is a persistent diagram (or transformations of it). Current research discusses the consequences of the many ways to exploit topological signatures, much less often the choice of the filtration, but to the best of our knowledge, the choice of the representation of a signal has not been the subject of any study yet. This paper attempts to provide some answers on the latter problem. To this end, we collected real audio data and built a comparative study to assess the quality of the discriminant information of the topological signatures extracted from three different representation spaces. Each audio signal is represented as i) an embedding of observed data in a higher dimensional space using Taken's representation, ii) a spectrogram viewed as a surface in a 3D ambient space, iii) the set of spectrogram's zeroes. From vowel audio recordings, we use topological signature for three prediction problems: speaker gender, vowel type, and individual. We show that topologically-augmented random forest improves the Out-of-Bag Error (OOB) over solely based Mel-Frequency Cepstral Coefficients (MFCC) for the last two problems. Our results also suggest that the topological information extracted from different signal representations is complementary, and that spectrogram's zeros offers the best improvement for gender prediction.

cs.SD

Detection and classification of vocal productions in large scale audio recordings

We propose an automatic data processing pipeline to extract vocal productions from large-scale natural audio recordings and classify these vocal productions. The pipeline is based on a deep neural network and adresses both issues simultaneously. Though a series of computationel steps (windowing, creation of a noise class, data augmentation, re-sampling, transfer learning, Bayesian optimisation), it automatically trains a neural network without requiring a large sample of labeled data and important computing resources. Our end-to-end methodology can handle noisy recordings made under different recording conditions. We test it on two different natural audio data sets, one from a group of Guinea baboons recorded from a primate research center and one from human babies recorded at home. The pipeline trains a model on 72 and 77 minutes of labeled audio recordings, with an accuracy of 94.58% and 99.76%. It is then used to process 443 and 174 hours of natural continuous recordings and it creates two new databases of 38.8 and 35.2 hours, respectively. We discuss the strengths and limitations of this approach that can be applied to any massive audio recording.

cs.SD

Tempered, Anti-trunctated, Multiple Importance Sampling

Importance sampling is a Monte Carlo method that introduces a proposal distribution to sample the space according to the target distribution. Yet calibration of the proposal distribution is essential to achieving efficiency, thus the resort to adaptive algorithms to tune this distribution. In the paper, we propose a new adpative importance sampling scheme, named Tempered Anti-truncated Adaptive Multiple Importance Sampling (TAMIS) algorithm. We combine a tempering scheme and a new nonlinear transformation of the weights we named anti-truncation. For efficiency, we were also concerned not to increase the number of evaluations of the target density. As a result, our proposal is an automatically tuned sequential algorithm that is robust to poor initial proposals, does not require gradient computations and scales well with the dimension.

stat.CO

Combined parameter and state inference with automatically calibrated ABC

State space models contain time-indexed parameters, termed states, as well as static parameters, simply termed parameters. The problem of inferring both static parameters as well as states simultaneously, based on time-indexed observations, is the subject of much recent literature. This problem is compounded once we consider models with intractable likelihoods. In these situations, some emerging approaches have incorporated existing likelihood-free techniques for static parameters, such as approximate Bayesian computation (ABC) into likelihood-based algorithms for combined inference of parameters and states. These emerging approaches currently require extensive manual calibration of a time-indexed tuning parameter: the acceptance threshold. We design an SMC$^2$ algorithm (Chopin et al., 2013, JRSS B) for likelihood-free approximation with automatically tuned thresholds. We prove consistency of the algorithm and discuss the proposed calibration. We demonstrate this algorithm's performance with three examples. We begin with two examples of state space models. The first example is a toy example, with an emission distribution that is a skew normal distribution. The second example is a stochastic volatility model involving an intractable stable distribution. The last example is the most challenging; it deals with an inhomogeneous Hawkes process.

stat.CO

ABC random forests for Bayesian parameter inference

This preprint has been reviewed and recommended by Peer Community In Evolutionary Biology (http://dx.doi.org/10.24072/pci.evolbiol.100036). Approximate Bayesian computation (ABC) has grown into a standard methodology that manages Bayesian inference for models associated with intractable likelihood functions. Most ABC implementations require the preliminary selection of a vector of informative statistics summarizing raw data. Furthermore, in almost all existing implementations, the tolerance level that separates acceptance from rejection of simulated parameter values needs to be calibrated. We propose to conduct likelihood-free Bayesian inferences about parameters with no prior selection of the relevant components of the summary statistics and bypassing the derivation of the associated tolerance level. The approach relies on the random forest methodology of Breiman (2001) applied in a (non parametric) regression setting. We advocate the derivation of a new random forest for each component of the parameter vector of interest. When compared with earlier ABC solutions, this method offers significant gains in terms of robustness to the choice of the summary statistics, does not depend on any type of tolerance level, and is a good trade-off in term of quality of point estimator precision and credible interval estimations for a given computing time. We illustrate the performance of our methodological proposal and compare it with earlier ABC methods on a Normal toy example and a population genetics example dealing with human population evolution. All methods designed here have been incorporated in the R package abcrf (version 1.7) available on CRAN.

stat.ME

Faster Hamiltonian Monte Carlo by Learning Leapfrog Scale: an offline randomized solution

We introduce a Hamiltonian Monte Carlo (HMC) methodology based on an offline empirical calibration of randomized leapfrog parameters. The approach, referred to as eHMC, where \textit{e} stands for empirical, leverages importance sampling to construct an empirical distribution on discretization parameters, thereby eliminating the need for manual burn-in diagnostics and online adaptation. The proposal distribution used in the calibration stage is obtained via a Population Monte Carlo scheme with tempering and relies on flexible parametric variational families such as normalizing flows. Once the calibration stage complete, the resulting algorithm defines a homogeneous Markov chain via a mixture of HMC kernels with a fixed mixing distribution, and hence preserves the target distribution. Numerical experiments indicate that eHMC can achieve competitive or improved sampling efficiency compared to the No-U-Turn Sampler (NUTS) in the case useful integration times can be summarized by the offline distribution. The comparison is assessed by standard efficiency metrics normalized by the number of leapfrog steps during the post-calibration sampling phase.

stat.CO

Bayesian functional linear regression with sparse step functions

The functional linear regression model is a common tool to determine the relationship between a scalar outcome and a functional predictor seen as a function of time. This paper focuses on the Bayesian estimation of the support of the coefficient function. To this aim we propose a parsimonious and adaptive decomposition of the coefficient function as a step function, and a model including a prior distribution that we name Bayesian functional Linear regression with Sparse Step functions (Bliss). The aim of the method is to recover areas of time which influences the most the outcome. A Bayes estimator of the support is built with a specific loss function, as well as two Bayes estimators of the coefficient function, a first one which is smooth and a second one which is a step function. The performance of the proposed methodology is analysed on various synthetic datasets and is illustrated on a black Périgord truffle dataset to study the influence of rainfall on the production.

stat.ME

Likelihood-free Model Choice

This document is an invited chapter covering the specificities of ABC model choice, intended for the incoming Handbook of ABC by Sisson, Fan, and Beaumont (2017). Beyond exposing the potential pitfalls of ABC based posterior probabilities, the review emphasizes mostly the solution proposed by Pudlo et al. (2016) on the use of random forests for aggregating summary statistics and and for estimating the posterior probability of the most likely model via a secondary random fores.

stat.ME

Resampling: an improvement of Importance Sampling in varying population size models

Sequential importance sampling algorithms have been defined to estimate likelihoods in models of ancestral population processes. However, these algorithms are based on features of the models with constant population size, and become inefficient when the population size varies in time, making likelihood-based inferences difficult in many demographic situations. In this work, we modify a previous sequential importance sampling algorithm to improve the efficiency of the likelihood estimation. Our procedure is still based on features of the model with constant size, but uses a resampling technique with a new resampling probability distribution depending on the pairwise composite likelihood. We tested our algorithm, called sequential importance sampling with resampling (SISR) on simulated data sets under different demographic cases. In most cases, we divided the computational cost by two for the same accuracy of inference, in some cases even by one hundred. This study provides the first assessment of the impact of such resampling techniques on parameter inference using sequential importance sampling, and extends the range of situations where likelihood inferences can be easily performed.

math.ST

Hidden Gibbs random fields model selection using Block Likelihood Information Criterion

Performing model selection between Gibbs random fields is a very challenging task. Indeed, due to the Markovian dependence structure, the normalizing constant of the fields cannot be computed using standard analytical or numerical methods. Furthermore, such unobserved fields cannot be integrated out and the likelihood evaluztion is a doubly intractable problem. This forms a central issue to pick the model that best fits an observed data. We introduce a new approximate version of the Bayesian Information Criterion. We partition the lattice into continuous rectangular blocks and we approximate the probability measure of the hidden Gibbs field by the product of some Gibbs distributions over the blocks. On that basis, we estimate the likelihood and derive the Block Likelihood Information Criterion (BLIC) that answers model choice questions such as the selection of the dependency structure or the number of latent states. We study the performances of BLIC for those questions. In addition, we present a comparison with ABC algorithms to point out that the novel criterion offers a better trade-off between time efficiency and reliable results.

stat.CO

Reliable ABC model choice via random forests

Approximate Bayesian computation (ABC) methods provide an elaborate approach to Bayesian inference on complex models, including model choice. Both theoretical arguments and simulation experiments indicate, however, that model posterior probabilities may be poorly evaluated by standard ABC techniques. We propose a novel approach based on a machine learning tool named random forests to conduct selection among the highly complex models covered by ABC algorithms. We thus modify the way Bayesian model selection is both understood and operated, in that we rephrase the inferential goal as a classification problem, first predicting the model that best fits the data with random forests and postponing the approximation of the posterior probability of the predicted MAP for a second stage also relying on random forests. Compared with earlier implementations of ABC model choice, the ABC random forest approach offers several potential improvements: (i) it often has a larger discriminative power among the competing models, (ii) it is more robust against the number and choice of statistics summarizing the data, (iii) the computing effort is drastically reduced (with a gain in computation efficiency of at least fifty), and (iv) it includes an approximation of the posterior probability of the selected model. The call to random forests will undoubtedly extend the range of size of datasets and complexity of models that ABC can handle. We illustrate the power of this novel methodology by analyzing controlled experiments as well as genuine population genetics datasets. The proposed methodologies are implemented in the R package abcrf available on the CRAN.

stat.ML

Adaptive ABC model choice and geometric summary statistics for hidden Gibbs random fields

Selecting between different dependency structures of hidden Markov random field can be very challenging, due to the intractable normalizing constant in the likelihood. We answer this question with approximate Bayesian computation (ABC) which provides a model choice method in the Bayesian paradigm. This comes after the work of Grelaud et al. (2009) who exhibited sufficient statistics on directly observed Gibbs random fields. But when the random field is latent, the sufficiency falls and we complement the set with geometric summary statistics. The general approach to construct these intuitive statistics relies on a clustering analysis of the sites based on the observed colors and plausible latent graphs. The efficiency of ABC model choice based on these statistics is evaluated via a local error rate which may be of independent interest. As a byproduct we derived an ABC algorithm that adapts the dimension of the summary statistics to the dataset without distorting the model selection.

math.ST

Consistency of the Adaptive Multiple Importance Sampling

Among Monte Carlo techniques, the importance sampling requires fine tuning of a proposal distribution, which is now fluently resolved through iterative schemes. The Adaptive Multiple Importance Sampling (AMIS) of Cornuet et al. (2012) provides a significant improvement in stability and effective sample size due to the introduction of a recycling procedure. However, the consistency of the AMIS estimator remains largely open. In this work we prove the convergence of the AMIS, at a cost of a slight modification in the learning process. Contrary to Douc et al. (2007a), results are obtained here in the asymptotic regime where the number of iterations is going to infinity while the number of drawings per iteration is a fixed, but growing sequence of integers. Hence some of the results shed new light on adaptive population Monte Carlo algorithms in that last regime.

stat.CO

Efficient learning in ABC algorithms

Approximate Bayesian Computation has been successfully used in population genetics to bypass the calculation of the likelihood. These methods provide accurate estimates of the posterior distribution by comparing the observed dataset to a sample of datasets simulated from the model. Although parallelization is easily achieved, computation times for ensuring a suitable approximation quality of the posterior distribution are still high. To alleviate the computational burden, we propose an adaptive, sequential algorithm that runs faster than other ABC algorithms but maintains accuracy of the approximation. This proposal relies on the sequential Monte Carlo sampler of Del Moral et al. (2012) but is calibrated to reduce the number of simulations from the model. The paper concludes with numerical experiments on a toy example and on a population genetic study of Apis mellifera, where our algorithm was shown to be faster than traditional ABC schemes.

stat.CO