Searcharxiv⌕ Search

arXiv subjects

Marc Hoffmann

Publications and source records attributed to Marc Hoffmann.

At least 37 records · Page 2Linked to original sources

Testing for high frequency features in a noisy signal

Given nonstationary data, one generally wants to extract the trend from the noise by smoothing or filtering. However, it is often important to delineate a third intermediate category, that we call high frequency (HF) features: this is the case in our motivating example, which consists in experimental measurements of the time-dynamics of depolymerising protein fibrils average size. One may intuitively visualise HF features as the presence of fast, possibly nonstationary and transient oscillations, distinct from a slowly-varying trend envelope. The aim of this article is to propose an empirical definition of HF features and construct estimators and statistical tests for their presence accordingly, when the data consists of a noisy nonstationary 1-dimensional signal. We propose a parametric characterization in the Fourier domain of the HF features by defining a maximal amplitude and distance to low frequencies of significant energy. We introduce a data-driven procedure to estimate these parameters, and compute a p-value proxy based on a statistical test for the presence of HF features. The test is first conducted on simulated signals where the ratio amplitude of the HF features to the level of the noise is controlled. The test detects HF features even when the level of noise is five times larger than the amplitude of the oscillations. In a second part, the test is conducted on experimental data from Prion disease experiments and it confirms the presence of HF features in these signals with significant confidence.

eess.SP↗

Efficient volatility estimation in a two-factor model

We statistically analyse a multivariate HJM diffusion model with stochastic volatility. The volatility process of the first factor is left totally unspecified while the volatility of the second factor is the product of an unknown process and an exponential function of time to maturity. This exponential term includes some real parameter measuring the rate of increase of the second factor as time goes to maturity. From historical data, we efficiently estimate the time to maturity parameter in the sense of constructing an estimator that achieves an optimal information bound in a semiparametric setting. We also identify nonparametrically the paths of the volatility processes and achieve minimax bounds. We address the problem of degeneracy that occurs when the dimension of the process is greater than two, and give in particular optimal limit theorems under suitable regularity assumptions on the drift process. We consistently analyse the numerical behaviour of our estimators on simulated and real datasets of prices of forward contracts on electricity markets.

math.ST↗

Nonparametric adaptive inference of birth and death models in a large population limit

Motivated by improving mortality tables from human demography databases, we investigate statistical inference of a stochastic age-evolving density of a population alimented by time inhomogeneous mortality and fertility. Asymptotics are taken as the size of the population grows within a limited time horizon: the observation gets closer to the solution of the Von Foerster Mc Kendrick equation, and the difficulty lies in controlling simultaneously the stochastic approximation to the limiting PDE in a suitable sense together with an appropriate parametrisation of the anisotropic solution. In this setting, we prove new concentration inequalities that enable us to implement the Goldenshluger-Lepski algorithm and derive oracle inequalities. We obtain minimax optimality and adaptation over a wide range of anisotropic Hölder smoothness classes.

math.ST↗

Statistical estimation in a randomly structured branching population

We consider a binary branching process structured by a stochastic trait that evolves according to a diffusion process that triggers the branching events, in the spirit of Kimmel's model of cell division with parasite infection. Based on the observation of the trait at birth of the first n generations of the process, we construct nonparametric estimator of the transition of the associated bifurcating chain and study the parametric estimation of the branching rate. In the limit, as n tends to infinity, we obtain asymptotic efficiency in the parametric case and minimax optimality in the nonparametric case.

math.ST↗

Early stopping for statistical inverse problems via truncated SVD estimation

We consider truncated SVD (or spectral cut-off, projection) estimators for a prototypical statistical inverse problem in dimension $D$. Since calculating the singular value decomposition (SVD) only for the largest singular values is much less costly than the full SVD, our aim is to select a data-driven truncation level $\widehat m\in\{1,\ldots,D\}$ only based on the knowledge of the first $\widehat m$ singular values and vectors. We analyse in detail whether sequential {\it early stopping} rules of this type can preserve statistical optimality. Information-constrained lower bounds and matching upper bounds for a residual based stopping rule are provided, which give a clear picture in which situation optimal sequential adaptation is feasible. Finally, a hybrid two-step approach is proposed which allows for classical oracle inequalities while considerably reducing numerical complexity.

math.ST↗

Optimal adaptation for early stopping in statistical inverse problems

For linear inverse problems $Y=\mathsf{A}μ+ξ$, it is classical to recover the unknown signal $μ$ by iterative regularisation methods $(\widehat μ^{(m)}, m=0,1,\ldots)$ and halt at a data-dependent iteration $τ$ using some stopping rule, typically based on a discrepancy principle, so that the weak (or prediction) squared-error $\|\mathsf{A}(\widehat μ^{(τ)}-μ)\|^2$ is controlled. In the context of statistical estimation with stochastic noise $ξ$, we study oracle adaptation (that is, compared to the best possible stopping iteration) in strong squared-error $E[\|\hat μ^{(τ)}-μ\|^2]$. For a residual-based stopping rule oracle adaptation bounds are established for general spectral regularisation methods. The proofs use bias and variance transfer techniques from weak prediction error to strong $L^2$-error, as well as convexity arguments and concentration bounds for the stochastic part. Adaptive early stopping for the Landweber method is studied in further detail and illustrated numerically.

math.ST↗

A recursive point process model for infectious diseases

We introduce a new type of point process model to describe the incidence of contagious diseases. The model is a variant of the Hawkes self-exciting process and exhibits similar clustering but without the restriction that the component describing the contagion must remain static over time. Instead, our proposed model prescribes that the degree of contagion (or productivity) changes as a function of the conditional intensity; of particular interest is the special case where the productivity is inversely proportional to the conditional intensity. The model incorporates the premise that when the disease occurs at very low frequency in the population, such as in the primary stages of an outbreak, then anyone with the disease is likely to have a high rate of transmission to others, whereas when the disease is prevalent in the population, then the transmission rate is lower due to human mitigation actions and prevention measures and a relatively high percentage of previous exposure in the total population. The model is said to be recursive, in the sense that the conditional intensity at any particular time depends on the productivity associated with previous points, and this productivity in turn depends on the conditional intensity at those points. Some basic properties of the model are derived, estimation and simulation are discussed, and the recursive model is shown to fit well to historic data on measles in Los Angeles, California, a relevant example given the 2017 outbreak of this disease in the same region.

stat.ME↗

On adaptive posterior concentration rates

We investigate the problem of deriving posterior concentration rates under different loss functions in nonparametric Bayes. We first provide a lower bound on posterior coverages of shrinking neighbourhoods that relates the metric or loss under which the shrinking neighbourhood is considered, and an intrinsic pre-metric linked to frequentist separation rates. In the Gaussian white noise model, we construct feasible priors based on a spike and slab procedure reminiscent of wavelet thresholding that achieve adaptive rates of contraction under $L^2$ or $L^{\infty}$ metrics when the underlying parameter belongs to a collection of Hölder balls and that moreover achieve our lower bound. We analyse the consequences in terms of asymptotic behaviour of posterior credible balls as well as frequentist minimax adaptive estimation. Our results are appended with an upper bound for the contraction rate under an arbitrary loss in a generic regular experiment. The upper bound is attained for certain sieve priors and enables to extend our results to density estimation.

math.ST↗

Nonparametric estimation of the division rate of an age dependent branching process

We study the nonparametric estimation of the branching rate $B(x)$ of a supercritical Bellman-Harris population: a particle with age $x$ has a random lifetime governed by $B(x)$; at its death time, it gives rise to $k \geq 2$ children with lifetimes governed by the same division rate and so on. We observe in continuous time the process over $[0,T]$. Asymptotics are taken as $T \rightarrow \infty$; the data are stochastically dependent and one has to face simultaneously censoring, bias selection and non-ancillarity of the number of observations. In this setting, under appropriate ergodicity properties, we construct a kernel-based estimator of $B(x)$ that achieves the rate of convergence $\exp(-λ_B \fracβ{2β+1}T)$, where $λ_B$ is the Malthus parameter and $β>0$ is the smoothness of the function $B(x)$ in a vicinity of $x$. We prove that this rate is optimal in a minimax sense and we relate it explicitly to classical nonparametric models such as density estimation observed on an appropriate (parameter dependent) scale. We also shed some light on the fact that estimation with kernel estimators based on data alive at time $T$ only is not sufficient to obtain optimal rates of convergence, a phenomenon which is specific to nonparametric estimation and that has been observed in other related growth-fragmentation models.

stat.OT↗

Adaptive estimation for bifurcating Markov chains

In a first part, we prove Bernstein-type deviation inequalities for bifurcating Markov chains (BMC) under a geometric ergodicity assumption, completing former results of Guyon and Bitseki Penda, Djellout and Guillin. These preliminary results are the key ingredient to implement nonparametric wavelet thresholding estimation procedures: in a second part, we construct nonparametric estimators of the transition density of a BMC, of its mean transition density and of the corresponding invariant density, and show smoothness adaptation over various multivariate Besov classes under $L^p$-loss error, for $1 \leq p < \infty$. We prove that our estimators are (nearly) optimal in a minimax sense. As an application, we obtain new results for the estimation of the splitting size-dependent rate of growth-fragmentation models and we extend the statistical study of bifurcating autoregressive processes.

math.ST↗

Statistical estimation of a growth-fragmentation model observed on a genealogical tree

We model the growth of a cell population by a piecewise deterministic Markov branching tree. Each cell splits into two offsprings at a division rate $B(x)$ that depends on its size $x$. The size of each cell grows exponentially in time, at a rate that varies for each individual. We show that the mean empirical measure of the model satisfies a growth-fragmentation type equation if structured in both size and growth rate as state variables. We construct a nonparametric estimator of the division rate $B(x)$ based on the observation of the population over different sampling schemes of size $n$ on the genealogical tree. Our estimator nearly achieves the rate $n^{-s/(2s+1)}$ in squared-loss error asymptotically. When the growth rate is assumed to be identical for every cell, we retrieve the classical growth-fragmentation model and our estimator improves on the rate $n^{-s/(2s+3)}$ obtained in \cite{DHRR, DPZ} through indirect observation schemes. Our method is consistently tested numerically and implemented on {\it Escherichia coli} data.

math.PR↗

High dimensional Hawkes processes

We generalise the construction of multivariate Hawkes processes to a possibly infinite network of counting processes on a directed graph $\mathbb G$. The process is constructed as the solution to a system of Poisson driven stochastic differential equations, for which we prove pathwise existence and uniqueness under some reasonable conditions. We next investigate how to approximate a standard $N$-dimensional Hawkes process by a simple inhomogeneous Poisson process in the mean-field framework where each pair of individuals interact in the same way, in the limit $N \rightarrow \infty$. In the so-called linear case for the interaction, we further investigate the large time behaviour of the process. We study in particular the stability of the central limit theorem when exchanging the limits $N, T\rightarrow \infty$ and exhibit different possible behaviours. We finally consider the case $\mathbb G = \mathbb Z^d$ with nearest neighbour interactions. In the linear case, we prove some (large time) laws of large numbers and exhibit different behaviours, reminiscent of the infinite setting. Finally we study the propagation of a {\it single impulsion} started at a given point of $\zz^d$ at time $0$. We compute the probability of extinction of such an impulsion and, in some particular cases, we can accurately describe how it propagates to the whole space.

math.PR↗

On adaptive inference and confidence bands

The problem of existence of adaptive confidence bands for an unknown density $f$ that belongs to a nested scale of Hölder classes over $\mathbb{R}$ or $[0,1]$ is considered. Whereas honest adaptive inference in this problem is impossible already for a pair of Hölder balls $Σ(r),Σ(s),r\ne s$, of fixed radius, a nonparametric distinguishability condition is introduced under which adaptive confidence bands can be shown to exist. It is further shown that this condition is necessary and sufficient for the existence of honest asymptotic confidence bands, and that it is strictly weaker than similar analytic conditions recently employed in Giné and Nickl [Ann. Statist. 38 (2010) 1122--1170]. The exceptional sets for which honest inference is not possible have vanishingly small probability under natural priors on Hölder balls $Σ(s)$. If no upper bound for the radius of the Hölder balls is known, a price for adaptation has to be paid, and near-optimal adaptation is possible for standard procedures. The implications of these findings for a general theory of adaptive inference are discussed.

math.ST↗

Scaling limits for Hawkes processes and application to financial statistics

We prove a law of large numbers and a functional central limit theorem for multivariate Hawkes processes observed over a time interval $[0,T]$ in the limit $T \rightarrow \infty$. We further exhibit the asymptotic behaviour of the covariation of the increments of the components of a multivariate Hawkes process, when the observations are imposed by a discrete scheme with mesh $Δ$ over $[0,T]$ up to some further time shift $τ$. The behaviour of this functional depends on the relative size of $Δ$ and $τ$ with respect to $T$ and enables to give a full account of the second-order structure. As an application, we develop our results in the context of financial statistics. We introduced in a previous work a microscopic stochastic model for the variations of a multivariate financial asset, based on Hawkes processes and that is confined to live on a tick grid. We derive and characterise the exact macroscopic diffusion limit of this model and show in particular its ability to reproduce important empirical stylised fact such as the Epps effect and the lead-lag effect. Moreover, our approach enable to track these effects across scales in rigorous mathematical terms.

math.PR↗

Adaptive wavelet estimation of the diffusion coefficient under additive error measurements

We study nonparametric estimation of the diffusion coefficient from discrete data, when the observations are blurred by additional noise. Such issues have been developed over the last 10 years in several application fields and in particular in high frequency financial data modelling, however mainly from a parametric and semiparametric point of view. This paper addresses the nonparametric estimation of the path of the (possibly stochastic) diffusion coefficient in a relatively general setting. By developing pre-averaging techniques combined with wavelet thresholding, we construct adaptive estimators that achieve a nearly optimal rate within a large scale of smoothness constraints of Besov type. Since the diffusion coefficient is usually genuinely random, we propose a new criterion to assess the quality of estimation; we retrieve the usual minimax theory when this approach is restricted to a deterministic diffusion coefficient. In particular, we take advantage of recent results of Reiss [33] of asymptotic equivalence between a Gaussian diffusion with additive noise and Gaussian white noise model, in order to prove a sharp lower bound.

math.ST↗

Statistical inference across time scales

We investigate statistical inference across time scales. We take as toy model the estimation of the intensity of a discretely observed compound Poisson process with symmetric Bernoulli jumps. We have data at different time scales: microscopic, intermediate and macroscopic. We quantify the smooth statistical transition from a microscopic Poissonian regime to a macroscopic Gaussian regime. The classical quadratic variation estimator is efficient in both microscopic and macroscopic scales but surprisingly shows a substantial loss of information in the intermediate scale that can be explicitly related to the sampling rate. We discuss the implications of these findings beyond this idealised framework.

math.ST↗

Nonparametric estimation of the division rate of a size-structured population

We consider the problem of estimating the division rate of a size-structured population in a nonparametric setting. The size of the system evolves according to a transport-fragmentation equation: each individual grows with a given transport rate, and splits into two offsprings of the same size, following a binary fragmentation process with unknown division rate that depends on its size. In contrast to a deterministic inverse problem approach, as in (Perthame, Zubelli, 2007) and (Doumic, Perthame, Zubelli, 2009), we take in this paper the perspective of statistical inference: our data consists in a large sample of the size of individuals, when the evolution of the system is close to its time-asymptotic behavior, so that it can be related to the eigenproblem of the considered transport-fragmentation equation (see \cite{PR} for instance). By estimating statistically each term of the eigenvalue problem and by suitably inverting a certain linear operator (see previously quoted articles), we are able to construct a more realistic estimator of the division rate that achieves the same optimal error bound as in related deterministic inverse problems. Our procedure relies on kernel methods with automatic bandwidth selection. It is inspired by model selection and recent results of Goldenschluger and Lepski.

math.ST↗

Statistical analysis of self-similar conservative fragmentation chains

We explore statistical inference in self-similar conservative fragmentation chains when only approximate observations of the sizes of the fragments below a given threshold are available. This framework, introduced by Bertoin and Martinez [Adv. Appl. Probab. 37 (2005) 553--570], is motivated by mineral crushing in the mining industry. The underlying object that can be identified from the data is the step distribution of the random walk associated with a randomly tagged fragment that evolves along the genealogical tree representation of the fragmentation process. We compute upper and lower rates of estimation in a parametric framework and show that in the nonparametric case, the difficulty of the estimation is comparable to ill-posed linear inverse problems of order 1 in signal denoising.

math.ST↗