SearcharxivSearch

arXiv subjects

Arnaud Guyader

Publications and source records attributed to Arnaud Guyader.

15 recordsLinked to original sources

On deviation probabilities in non-parametric regression with heavy-tailed noise

This paper is devoted to the problem of determining the concentration bounds that are achievable in non-parametric regression. We consider the setting where features are supported on a bounded subset of $\mathbb{R}^d$, the regression function is Lipschitz, and the noise is only assumed to have a finite second moment. We first specify the fundamental limits of the problem by establishing a general lower bound on deviation probabilities, and then construct explicit estimators that achieve this bound. These estimators are obtained by applying the median-of-means principle to classical local averaging rules in non-parametric regression, including nearest neighbors and kernel procedures.

math.ST

On the Hill relation and the mean reaction time for metastable processes

We illustrate how the Hill relation and the notion of quasi-stationary distribution can be used to analyse the biasing error introduced by many numerical procedures that have been proposed in the literature, in particular in molecular dynamics, to compute mean reaction times between metastable states for Markov processes. The theoretical findings are illustrated on various examples demonstrating the sharpness of the biasing error analysis as well as the applicability of our study to elliptic diffusions.

math.PR

Efficient large deviation estimation based on importance sampling

We present a complete framework for determining the asymptotic (or logarithmic) efficiency of estimators of large deviation probabilities and rate functions based on importance sampling. The framework relies on the idea that importance sampling in that context is fully characterized by the joint large deviations of two random variables: the observable defining the large deviation probability of interest and the likelihood factor (or Radon-Nikodym derivative) connecting the original process and the modified process used in importance sampling. We recover with this framework known results about the asymptotic efficiency of the exponential tilting and obtain new necessary and sufficient conditions for a general change of process to be asymptotically efficient. This allows us to construct new examples of efficient estimators for sample means of random variables that do not have the exponential tilting form. Other examples involving Markov chains and diffusions are presented to illustrate our results.

cond-mat.stat-mech

Recursive Estimation of a Failure Probability for a Lipschitz Function

Let g : $Ω$ = [0, 1] d $\rightarrow$ R denote a Lipschitz function that can be evaluated at each point, but at the price of a heavy computational time. Let X stand for a random variable with values in $Ω$ such that one is able to simulate, at least approximately, according to the restriction of the law of X to any subset of $Ω$. For example, thanks to Markov chain Monte Carlo techniques, this is always possible when X admits a density that is known up to a normalizing constant. In this context, given a deterministic threshold T such that the failure probability p := P(g(X) > T) may be very low, our goal is to estimate the latter with a minimal number of calls to g. In this aim, building on Cohen et al. [9], we propose a recursive and optimal algorithm that selects on the fly areas of interest and estimate their respective probabilities.

math.PR

Variance Estimation in Adaptive Sequential Monte Carlo

Sequential Monte Carlo (SMC) methods represent a classical set of techniques to simulate a sequence of probability measures through a simple selection/mutation mechanism. However, the associated selection functions and mutation kernels usually depend on tuning parameters that are of first importance for the efficiency of the algorithm. A standard way to address this problem is to apply Adaptive Sequential Monte Carlo (ASMC) methods, which consist in exploiting the information given by the history of the sample to tune the parameters. This article is concerned with variance estimation in such ASMC methods. Specifically, we focus on the case where the asymptotic variance coincides with the one of the "limiting" Sequential Monte Carlo algorithm as defined by Beskos et al. (2016). We prove that, under natural assumptions, the estimator introduced by Lee and Whiteley (2018) in the nonadaptive case (i.e., SMC) is also a consistent estimator of the asymptotic variance for ASMC methods. To do this, we introduce a new estimator that is expressed in terms of coalescent tree-based measures, and explain its connection with the previous one. Our estimator is constructed by tracing the genealogy of the associated Interacting Particle System. The tools we use connect the study of Particle Markov Chain Monte Carlo methods and the variance estimation problem in SMC methods. As such, they may give some new insights when dealing with complex genealogy-involved problems of Interacting Particle Systems in more general scenarios.

math.ST

On Synchronized Fleming-Viot Particle Systems

This article presents a variant of Fleming-Viot particle systems, which are a standard way to approximate the law of a Markov process with killing as well as related quantities. Classical Fleming-Viot particle systems proceed by simulating $N$ trajectories, or particles, according to the dynamics of the underlying process, until one of them is killed. At this killing time, the particle is instantaneously branched on one of the $(N-1)$ other ones, and so on until a fixed and finite final time $T$. In our variant, we propose to wait until $K$ particles are killed and then rebranch them independently on the $(N-K)$ alive ones. Specifically, we focus our attention on the large population limit and the regime where $K/N$ has a given limit when $N$ goes to infinity. In this context, we establish consistency and asymptotic normality results. The variant we propose is motivated by applications in rare event estimation problems.

math.PR

On the Asymptotic Normality of Adaptive Multilevel Splitting

Adaptive Multilevel Splitting (AMS for short) is a generic Monte Carlo method for Markov processes that simulates rare events and estimates associated probabilities. Despite its practical efficiency, there are almost no theoretical results on the convergence of this algorithm. The purpose of this paper is to prove both consistency and asymptotic normality results in a general setting. This is done by associating to the original Markov process a level-indexed process, also called a stochastic wave, and by showing that AMS can then be seen as a Fleming-Viot type particle system. This being done, we can finally apply general results on Fleming-Viot particle systems that we have recently obtained.

math.PR

A Central Limit Theorem for Fleming-Viot Particle Systems with Soft Killing

The distribution of a Markov process with killing, conditioned to be still alive at a given time, can be approximated by a Fleming-Viot type particle system. In such a system, each particle is simulated independently according to the law of the underlying Markov process, and branches onto another particle at each killing time. The consistency of this method in the large population limit was the subject of several recent articles. In the present paper, we go one step forward and prove a central limit theorem for the law of the Fleming-Viot particle system at a given time under two conditions: a "soft killing" assumption and a boundedness condition involving the "carré du champ" operator of the underlying Markov process.

math.PR

A Central Limit Theorem for Fleming-Viot Particle Systems with Hard Killing

Fleming-Viot type particle systems represent a classical way to approximate the distribution of a Markov process with killing, given that it is still alive at a final deterministic time. In this context, each particle evolves independently according to the law of the underlying Markov process until its killing, and then branches instantaneously on another randomly chosen particle. While the consistency of this algorithm in the large population limit has been recently studied in several articles, our purpose here is to prove Central Limit Theorems under very general assumptions. For this, we only suppose that the particle system does not explode in finite time, and that the jump and killing times have atomless distributions. In particular, this includes the case of elliptic diffusions with hard killing.

math.PR

Fluctuation Analysis of Adaptive Multilevel Splitting

Multilevel Splitting is a Sequential Monte Carlo method to simulate realisations of a rare event as well as to estimate its probability. This article is concerned with the convergence and the fluctuation analysis of Adaptive Multilevel Splitting techniques. In contrast to their fixed level version, adaptive techniques estimate the sequence of levels on the fly and in an optimal way, with only a low additional computational cost. However, very few convergence results are available for this class of adaptive branching models, mainly because the sequence of levels depends on the occupation measures of the particle systems. This article proves the consistency of these methods as well as a central limit theorem. In particular, we show that the precision of the adaptive version is the same as the one of the fixed-levels version where the levels would have been placed in an optimal manner.

math.ST

New Insights Into Approximate Bayesian Computation

Approximate Bayesian Computation (ABC for short) is a family of computational techniques which offer an almost automated solution in situations where evaluation of the posterior likelihood is computationally prohibitive, or whenever suitable likelihoods are not available. In the present paper, we analyze the procedure from the point of view of k-nearest neighbor theory and explore the statistical properties of its outputs. We discuss in particular some asymptotic features of the genuine conditional density estimate associated with ABC, which is an interesting hybrid between a k-nearest neighbor and a kernel method.

math.ST

Iterative Isotonic Regression

This article introduces a new nonparametric method for estimating a univariate regression function of bounded variation. The method exploits the Jordan decomposition which states that a function of bounded variation can be decomposed as the sum of a non-decreasing function and a non-increasing function. This suggests combining the backfitting algorithm for estimating additive functions with isotonic regression for estimating monotone functions. The resulting iterative algorithm is called Iterative Isotonic Regression (I.I.R.). The main technical result in this paper is the consistency of the proposed estimator when the number of iterations $k_n$ grows appropriately with the sample size $n$. The proof requires two auxiliary results that are of interest in and by themselves: firstly, we generalize the well-known consistency property of isotonic regression to the framework of a non-monotone regression function, and secondly, we relate the backfitting algorithm to Von Neumann's algorithm in convex analysis.

math.ST

A Geometrical Approach to Iterative Isotone Regression

In the present paper, we propose and analyze a novel method for estimating a univariate regression function of bounded variation. The underpinning idea is to combine two classical tools in nonparametric statistics, namely isotonic regression and the estimation of additive models. A geometrical interpretation enables us to link this iterative method with Von Neumann's algorithm. Moreover, making a connection with the general property of isotonicity of projection onto convex cones, we derive another equivalent algorithm and go further in the analysis. As iterating the algorithm leads to overfitting, several practical stopping criteria are also presented and discussed.

math.ST

On the length of one-dimensional reactive paths

Motivated by some numerical observations on molecular dynamics simulations, we analyze metastable trajectories in a very simplecsetting, namely paths generated by a one-dimensional overdamped Langevin equation for a double well potential. More precisely, we are interested in so-called reactive paths, namely trajectories which leave definitely one well and reach the other one. The aim of this paper is to precisely analyze the distribution of the lengths of reactive paths in the limit of small temperature, and to compare the theoretical results to numerical results obtained by a Monte Carlo method, namely the multi-level splitting approach.

math.PR

A multiple replica approach to simulate reactive trajectories

A method to generate reactive trajectories, namely equilibrium trajectories leaving a metastable state and ending in another one is proposed. The algorithm is based on simulating in parallel many copies of the system, and selecting the replicas which have reached the highest values along a chosen one-dimensional reaction coordinate. This reaction coordinate does not need to precisely describe all the metastabilities of the system for the method to give reliable results. An extension of the algorithm to compute transition times from one metastable state to another one is also presented. We demonstrate the interest of the method on two simple cases: a one-dimensional two-well potential and a two-dimensional potential exhibiting two channels to pass from one metastable state to another one.

cond-mat.stat-mech