Searcharxiv⌕ Search

arXiv subjects

Arnaud Gloter

Publications and source records attributed to Arnaud Gloter.

At least 19 recordsLinked to original sources

Factorization by extremal privacy mechanisms: new insights into efficiency

We study the problem of efficiency under $α$ local differential privacy ($α$ LDP) in both discrete and continuous settings. Building on a factorization lemma, which shows that any privacy mechanism can be decomposed into an extremal mechanism followed by additional randomization, we reduce the Fisher information maximization problem to a search over extremal mechanisms. The representation of extremal mechanisms requires working in infinite dimensional spaces and invokes advanced tools from convex and functional analysis, such as Choquet's theorem. Our analysis establishes matching upper and lower bounds on the Fisher information in the high privacy regime ($α\to 0$), and proves that the maximization problem always admits a solution for any $α$. As a concrete application, we consider the problem of estimating the parameter of a uniform distribution on $[0, θ]$ under $α$ LDP. Guided by our theoretical findings, we design an extremal mechanism that yields a consistent and asymptotically efficient estimator in high privacy regime. Numerical experiments confirm our theoretical results.

math.ST↗

Risk comparison theorems and application to deep learning of diffusion coefficients

We investigate the nonparametric estimation of the diffusion matrix in stochastic differential equations featuring multidimensional, strong mixing covariate processes. We propose a flexible statistical framework based on general function classes that does not require a linear basis representation, rendering our results directly applicable to deep neural network estimators. Our approach employs a two-step estimation procedure: constructing a preliminary nonparametric quasi-likelihood estimator and subsequently regularizing it via a $β$-Hölder class approximation. We establish general risk comparison theorems between empirical and generalization risks for arbitrary estimators without relying on a specific probabilistic structure of the underlying process. In diffusion matrix learning based on $n + 1$ observations over the time interval $[0,T]$, the derived upper bounds capture the intrinsic interplay between the $T$-rate, associated with the mixing behavior of the covariate process, and the intrinsic $n$-rate governing the volatility estimation.

math.ST↗

Drift estimation for rough processes under small noise asymptotic : trajectory fitting method

We consider a process $X^\ve$ that solves a stochastic Volterra equation with an unknown parameter $θ^\star$ in the drift function. The Volterra kernel is singular, and includes as an example, $K\_0(u)=c u^{α-1/2} \id{u>0}$ with $α\in (0,1/2)$. It is assumed that the diffusion coefficient is proportional to $\ve \to 0$. From an observation of the path $(X^\ve\_s)\_{s\in[0,T]}$, we construct a Trajectory Fitting Estimator, which is shown to be consistent and asymptotically normal. We also specify identifiability conditions insuring the $L^p$ convergence of the estimator.

math.ST↗

Drift estimation for rough processes under small noise asymptotic : QMLE approach

We consider a process $X^\ve$ solution of a stochastic Volterra equation with an unknown parameter $θ^\star$ in the drift function. The Volterra kernel is singular near zero, exhibiting a behavior comparable to $K\_0(u)=cu^{α-1} \id{u>0}$ with $α\in (1/2,1)$.It is assumed that the diffusion coefficient is proportional to $\ve \to 0$. Based on discrete observations, with a mesh size $h\to0$, of the Volterra process, we construct a Quasi Maximum Likelihood Estimator. The main step is to assess the error arising in the reconstruction of the path of a semimartingale from the inversion of the Volterra kernel. We show that this error decreases as $h^{1/2}$ regardless of the value of $α$. Then, we can introduce an explicit contrast function, which yields an efficient estimator when $\ve \to 0$.

math.ST↗

On the role of symmetry for staircase mechanisms in local differential privacy efficiency across different privacy regimes

We investigate the structural foundations of statistical efficiency under $α$-local differential privacy, with a focus on maximizing Fisher information. Building on the role of continuous staircase mechanisms, we identify a fundamental symmetry regarding the extremal values $1$ and $e^α$. We demonstrate that when the optimal measure satisfies this symmetry, the Fisher information admits a closed-form expression. More generally, we derive a decomposition of the Fisher information into symmetric and asymmetric components, scaling as $α^{2}$ and $α^{3}$, respectively, for $α\to 0$. This reveals that, if in the high-privacy regime asymmetry is negligible, it is no longer the case as privacy constraints are relaxed. Motivated by this, we introduce a class of fully asymmetric privacy mechanisms constructed via pushforward mappings, proving that-unlike their symmetric counterparts-they recover the full Fisher information of the non-private model as $α\to \infty$. We bridge the gap between theory and practice by providing a tractable implementation of these mechanisms, governed by a tuning parameter $c$. This parameter allows for a smooth interpolation between the symmetric regime and the fully asymmetric regime. Furthermore, we demonstrate the versatility of this framework by showing that it encompasses the binomial mechanism as a limiting case.

math.ST↗

Nonparametric estimation of the stationary density for Hawkes-diffusion systems with known and unknown intensity

We investigate the nonparametric estimation problem of the density $π$, representing the stationary distribution of a two-dimensional system $\left(Z_t\right)_{t \in[0, T]}=\left(X_t, λ_t\right)_{t \in[0, T]}$. In this system, $X$ is a Hawkes-diffusion process, and $λ$ denotes the stochastic intensity of the Hawkes process driving the jumps of $X$. Based on the continuous observation of a path of $(X_t)$ over $[0, T]$, and initially assuming that $λ$ is known, we establish the convergence rate of a kernel estimator $\widehatπ\left(x^*, y^*\right)$ of $π\left(x^*,y^*\right)$ as $T \rightarrow \infty$. Interestingly, this rate depends on the value of $y^*$ influenced by the baseline parameter of the Hawkes intensity process. From the rate of convergence of $\widehatπ\left(x^*,y^*\right)$, we derive the rate of convergence for an estimator of the invariant density $λ$. Subsequently, we extend the study to the case where $λ$ is unknown, plugging an estimator of $λ$ in the kernel estimator and deducing new rates of convergence for the obtained estimator. The proofs establishing these convergence rates rely on probabilistic results that may hold independent interest. We introduce a Girsanov change of measure to transform the Hawkes process with intensity $λ$ into a Poisson process with constant intensity. To achieve this, we extend a bound for the exponential moments for the Hawkes process, originally established in the stationary case, to the non-stationary case. Lastly, we conduct a numerical study to illustrate the obtained rates of convergence of our estimators.

math.ST↗

Minimax rate for multivariate data under componentwise local differential privacy constraints

Our research delves into the balance between maintaining privacy and preserving statistical accuracy when dealing with multivariate data that is subject to \textit{componentwise local differential privacy} (CLDP). With CLDP, each component of the private data is made public through a separate privacy channel. This allows for varying levels of privacy protection for different components or for the privatization of each component by different entities, each with their own distinct privacy policies. We develop general techniques for establishing minimax bounds that shed light on the statistical cost of privacy in this context, as a function of the privacy levels $α_1, ... , α_d$ of the $d$ components. We demonstrate the versatility and efficiency of these techniques by presenting various statistical applications. Specifically, we examine nonparametric density and covariance estimation under CLDP, providing upper and lower bounds that match up to constant factors, as well as an associated data-driven adaptive procedure. Furthermore, we quantify the probability of extracting sensitive information from one component by exploiting the fact that, on another component which may be correlated with the first, a smaller degree of privacy protection is guaranteed.

math.ST↗

Evolving privacy: drift parameter estimation for discretely observed i.i.d. diffusion processes under LDP

The problem of estimating a parameter in the drift coefficient is addressed for $N$ discretely observed independent and identically distributed stochastic differential equations (SDEs). This is done considering additional constraints, wherein only public data can be published and used for inference. The concept of local differential privacy (LDP) is formally introduced for a system of stochastic differential equations. The objective is to estimate the drift parameter by proposing a contrast function based on a pseudo-likelihood approach. A suitably scaled Laplace noise is incorporated to meet the privacy requirements. Our key findings encompass the derivation of explicit conditions tied to the privacy level. Under these conditions, we establish the consistency and asymptotic normality of the associated estimator. Notably, the convergence rate is intricately linked to the privacy level, and is some situations may be completely different from the case where privacy constraints are ignored. Our results hold true as the discretization step approaches zero and the number of processes $N$ tends to infinity.

math.ST↗

Malliavin calculus for the optimal estimation of the invariant density of discretely observed diffusions in intermediate regime

Let $(X_t)_{t \ge 0}$ be solution of a one-dimensional stochastic differential equation. Our aim is to study the convergence rate for the estimation of the invariant density in intermediate regime, assuming that a discrete observation of the process $(X_t)_{t \in [0, T]}$ is available, when $T$ tends to $\infty$. We find the convergence rates associated to the kernel density estimator we proposed and a condition on the discretization step $Δ_n$ which plays the role of threshold between the intermediate regime and the continuous case. In intermediate regime the convergence rate is $n^{- \frac{2 β}{2 β+ 1}}$, where $β$ is the smoothness of the invariant density. After that, we complement the upper bounds previously found with a lower bound over the set of all the possible estimator, which provides the same convergence rate: it means it is not possible to propose a different estimator which achieves better convergence rates. This is obtained by the two hypotheses method; the most challenging part consists in bounding the Hellinger distance between the laws of the two models. The key point is a Malliavin representation for a score function, which allows us to bound the Hellinger distance through a quantity depending on the Malliavin weight.

math.ST↗

Minimax rate of estimation for invariant densities associated to continuous stochastic differential equations over anisotropic Holder classes

We study the problem of the nonparametric estimation for the density $π$ of the stationary distribution of a $d$-dimensional stochastic differential equation $(X_t)_{t \in [0, T]}$. From the continuous observation of the sampling path on $[0, T]$, we study the estimation of $π(x)$ as $T$ goes to infinity. For $d\ge2$, we characterize the minimax rate for the $\mathbf{L}^2$-risk in pointwise estimation over a class of anisotropic Hölder functions $π$ with regularity $β= (β_1, ... , β_d)$. For $d \ge 3$, our finding is that, having ordered the smoothness such that $β_1 \le ... \le β_d$, the minimax rate depends on whether $β_2 < β_3$ or $β_2 = β_3$. In the first case, this rate is $(\frac{\log T}{T})^γ$, and in the second case, it is $(\frac{1}{T})^γ$, where $γ$ is an explicit exponent dependent on the dimension and $\barβ_3$, the harmonic mean of smoothness over the $d$ directions after excluding $β_1$ and $β_2$, the smallest ones. We also demonstrate that kernel-based estimators achieve the optimal minimax rate. Furthermore, we propose an adaptive procedure for both $L^2$ integrated and pointwise risk. In the two-dimensional case, we show that kernel density estimators achieve the rate $\frac{\log T}{T}$, which is optimal in the minimax sense. Finally we illustrate the validity of our theoretical findings by proposing numerical results.

math.ST↗

Estimation of the invariant density for discretely observed diffusion processes: impact of the sampling and of the asynchronicity

We aim at estimating in a non-parametric way the density $π$ of the stationary distribution of a $d$-dimensional stochastic differential equation $(X_t)_{t \in [0, T]}$, for $d \ge 2$, from the discrete observations of a finite sample $X_{t_0}$, ... , $X_{t_n}$ with $0= t_0 < t_1 < ... < t_n =: T_n$. We propose a kernel density estimator and we study its convergence rates for the pointwise estimation of the invariant density under anisotropic Hölder smoothness constraints. First of all, we find some conditions on the discretization step that ensures it is possible to recover the same rates as if the continuous trajectory of the process was available. Such rates are optimal and new in the context of density estimator. Then we deal with the case where such a condition on the discretization step is not satisfied, which we refer to as intermediate regime. In this new regime we identify the convergence rate for the estimation of the invariant density over anisotropic Hölder classes, which is the same convergence rate as for the estimation of a probability density belonging to an anisotropic Hölder class, associated to $n$ iid random variables $X_1, ..., X_n$. After that we focus on the asynchronous case, in which each component can be observed at different time points. Even if the asynchronicity of the observations complexifies the computation of the variance of the estimator, we are able to find conditions ensuring that this variance is comparable to the one of the continuous case. We also exhibit that the non synchronicity of the data introduces additional bias terms in the study of the estimator.

math.ST↗

On the nonparametric inference of coefficients of self-exciting jump-diffusion

In this paper, we consider a one-dimensional diffusion process with jumps driven by a Hawkes process. We are interested in the estimations of the volatility function and of the jump function from discrete high-frequency observations in a long time horizon which remained an open question until now. First, we propose to estimate the volatility coefficient. For that, we introduce a truncation function in our estimation procedure that allows us to take into account the jumps of the process and estimate the volatility function on a linear subspace of L2(A) where A is a compact interval of R. We obtain a bound for the empirical risk of the volatility estimator, ensuring its consistency, and then we study an adaptive estimator w.r.t. the regularity. Then, we define an estimator of a sum between the volatility and the jump coefficient modified with the conditional expectation of the intensity of the jumps. We also establish a bound for the empirical risk for the non-adaptive estimators of this sum, the convergence rate up to the regularity of the true function, and an oracle inequality for the final adaptive estimator.Finally, we give a methodology to recover the jump function in some applications. We conduct a simulation study to measure our estimators' accuracy in practice and discuss the possibility of recovering the jump function from our estimation procedure.

math.ST↗

Joint estimation for volatility and drift parameters of ergodic jump diffusion processes via contrast function

In this paper we consider an ergodic diffusion process with jumps whose drift coefficient depends on $μ$ and volatility coefficient depends on $σ$, two unknown parameters. We suppose that the process is discretely observed at the instants (t n i)i=0,...,n with $Δ$n = sup i=0,...,n--1 (t n i+1 -- t n i) $\rightarrow$ 0. We introduce an estimator of $θ$ := ($μ$, $σ$), based on a contrast function, which is asymptotically gaussian without requiring any conditions on the rate at which $Δ$n $\rightarrow$ 0, assuming a finite jump activity. This extends earlier results where a condition on the step discretization was needed (see [13],[28]) or where only the estimation of the drift parameter was considered (see [2]). In general situations, our contrast function is not explicit and in practise one has to resort to some approximation. We propose explicit approximations of the contrast function, such that the estimation of $θ$ is feasible under the condition that n$Δ$ k n $\rightarrow$ 0 where k > 0 can be arbitrarily large. This extends the results obtained by Kessler [17] in the case of continuous processes. Efficient drift estimation, efficient volatility estimation,ergodic properties, high frequency data, L{é}vy-driven SDE, thresholding methods.

math.ST↗

Unbiased truncated quadratic variation for volatility estimation in jump diffusion processes

The problem of integrated volatility estimation for the solution X of a stochastic differential equation with L{é}vy-type jumps is considered under discrete high-frequency observations in both short and long time horizon. We provide an asymptotic expansion for the integrated volatility that gives us, in detail, the contribution deriving from the jump part. The knowledge of such a contribution allows us to build an unbiased version of the truncated quadratic variation, in which the bias is visibly reduced. In earlier results the condition $β$ > 1 2(2--$α$) on $β$ (that is such that (1/n) $β$ is the threshold of the truncated quadratic variation) and on the degree of jump activity $α$ was needed to have the original truncated realized volatility well-performed (see [22], [13]). In this paper we theoretically relax this condition and we show that our unbiased estimator achieves excellent numerical results for any couple ($α$, $β$). L{é}vy-driven SDE, integrated variance, threshold estimator, convergence speed, high frequency data.

math.ST↗

Adaptive and non-adaptive estimation for degenerate diffusion processes

We discuss parametric estimation of a degenerate diffusion system from time-discrete observations. The first component of the degenerate diffusion system has a parameter $θ_1$ in a non-degenerate diffusion coefficient and a parameter $θ_2$ in the drift term. The second component has a drift term parameterized by $θ_3$ and no diffusion term. Asymptotic normality is proved in three different situations for an adaptive estimator for $θ_3$ with some initial estimators for ($θ_1$ , $θ_2$), an adaptive one-step estimator for ($θ_1$ , $θ_2$ , $θ_3$) with some initial estimators for them, and a joint quasi-maximum likelihood estimator for ($θ_1$ , $θ_2$ , $θ_3$) without any initial estimator. Our estimators incorporate information of the increments of both components. Thanks to this construction, the asymptotic variance of the estimators for $θ_1$ is smaller than the standard one based only on the first component. The convergence of the estimators for $θ_3$ is much faster than the other parameters. The resulting asymptotic variance is smaller than that of an estimator only using the increments of the second component.

math.ST↗

Rate of Estimation for the Stationary Distribution of Stochastic Damping Hamiltonian Systems with Continuous Observations

We study the problem of the non-parametric estimation for the density $π$ of the stationary distribution of a stochastic two-dimensional damping Hamiltonian system $(Z_t)_{t\in[0,T]}=(X_t,Y_t)_{t \in [0,T]}$. From the continuous observation of the sampling path on $[0,T]$, we study the rate of estimation for $π(x_0,y_0)$ as $T \to \infty$. We show that kernel based estimators can achieve the rate $T^{-v}$ for some explicit exponent $v \in (0,1/2)$. One finding is that the rate of estimation depends on the smoothness of $π$ and is completely different with the rate appearing in the standard i.i.d.\ setting or in the case of two-dimensional non degenerate diffusion processes. Especially, this rate depends also on $y_0$. Moreover, we obtain a minimax lower bound on the $L^2$-risk for pointwise estimation, with the same rate $T^{-v}$, up to $\log(T)$ terms.

math.ST↗

Invariant density adaptive estimation for ergodic jump diffusion processes over anisotropic classes

We consider the solution X = (Xt) t$\ge$0 of a multivariate stochastic differential equation with Levy-type jumps and with unique invariant probability measure with density $μ$. We assume that a continuous record of observations X T = (Xt) 0$\le$t$\le$T is available. In the case without jumps, Reiss and Dalalyan (2007) and Strauch (2018) have found convergence rates of invariant density estimators, under respectively isotropic and anisotropic H{ö}lder smoothness constraints, which are considerably faster than those known from standard multivariate density estimation. We extend the previous works by obtaining, in presence of jumps, some estimators which have the same convergence rates they had in the case without jumps for d $\ge$ 2 and a rate which depends on the degree of the jumps in the one-dimensional setting. We propose moreover a data driven bandwidth selection procedure based on the Goldensh-luger and Lepski (2011) method which leads us to an adaptive non-parametric kernel estimator of the stationary density $μ$ of the jump diffusion X. Adaptive bandwidth selection, anisotropic density estimation, ergodic diffusion with jumps, L{é}vy driven SDE

math.ST↗