SearcharxivSearch

arXiv subjects

Dominique Picard

Publications and source records attributed to Dominique Picard.

At least 19 recordsLinked to original sources

Uncovering differential equations from data with hidden variables

SINDy is a method for learning system of differential equations from data by solving a sparse linear regression optimization problem [Brunton et al., 2016]. In this article, we propose an extension of the SINDy method that learns systems of differential equations in cases where some of the variables are not observed. Our extension is based on regressing a higher order time derivative of a target variable onto a dictionary of functions that includes lower order time derivatives of the target variable. We evaluate our method by measuring the prediction accuracy of the learned dynamical systems on synthetic data and on a real data-set of temperature time series provided by the Réseau de Transport d'Électricité (RTE). Our method provides high quality short-term forecasts and it is orders of magnitude faster than competing methods for learning differential equations with latent variables.

stat.ML

Clustering high dimensional meteorological scenarios: results and performance index

The Reseau de Transport d'Electricité (RTE) is the French main electricity network operational manager and dedicates large number of resources and efforts towards understanding climate time series data. We discuss here the problem and the methodology of grouping and selecting representatives of possible climate scenarios among a large number of climate simulations provided by RTE. The data used is composed of temperature times series for 200 different possible scenarios on a grid of geographical locations in France. These should be clustered in order to detect common patterns regarding temperatures curves and help to choose representative scenarios for network simulations, which in turn can be used for energy optimisation. We first show that the choice of the distance used for the clustering has a strong impact on the meaning of the results: depending on the type of distance used, either spatial or temporal patterns prevail. Then we discuss the difficulty of fine-tuning the distance choice (combined with a dimension reduction procedure) and we propose a methodology based on a carefully designed index.

stat.AP

Convergence rates for smooth k-means change-point detection

In this paper, we consider the estimation of a change-point for possibly high-dimensional data in a Gaussian model, using a k-means method. We prove that, up to a logarithmic term, this change-point estimator has a minimax rate of convergence. Then, considering the case of sparse data, with a Sobolev regularity, we propose a smoothing procedure based on Lepski's method and show that the resulting estimator attains the optimal rate of convergence. Our results are illustrated by some simulations. As the theoretical statement relying on Lepski's method depends on some unknown constant, practical strategies are suggested to perform an optimal smoothing.

math.ST

Statistical learning for wind power : a modeling and stability study towards forecasting

We focus on wind power modeling using machine learning techniques. We show on real data provided by the wind energy company Ma{ï}a Eolis, that parametric models, even following closely the physical equation relating wind production to wind speed are outperformed by intelligent learning algorithms. In particular, the CART-Bagging algorithm gives very stable and promising results. Besides, as a step towards forecast, we quantify the impact of using deteriorated wind measures on the performances. We show also on this application that the default methodology to select a subset of predictors provided in the standard random forest package can be refined, especially when there exists among the predictors one variable which has a major impact.

stat.AP

Regularity of Gaussian Processes on Dirichlet spaces

We are interested in the regularity of centered Gaussian processes (Z_x), x in M indexed by compact metric spaces M. It is shown that the almost everywhere Besov space regularity of such a process is (almost) equivalent to the Besov regularity of the covariance K(x,y) = E(Z_xZ_y) under the assumption that (i) there is an underlying Dirichlet structure on M which determines the Besov space regularity, and (ii) the operator K with kernel K(x,y) and the underlying operator A of the Dirichlet structure commute. As an application of this result we establish the Besov regularity of Gaussian processes indexed by compact homogeneous spaces and, in particular, by the sphere.

math.PR

Testing the isotropy of high energy cosmic rays using spherical needlets

For many decades, ultrahigh energy charged particles of unknown origin that can be observed from the ground have been a puzzle for particle physicists and astrophysicists. As an attempt to discriminate among several possible production scenarios, astrophysicists try to test the statistical isotropy of the directions of arrival of these cosmic rays. At the highest energies, they are supposed to point toward their sources with good accuracy. However, the observations are so rare that testing the distribution of such samples of directional data on the sphere is nontrivial. In this paper, we choose a nonparametric framework that makes weak hypotheses on the alternative distributions and allows in turn to detect various and possibly unexpected forms of anisotropy. We explore two particular procedures. Both are derived from fitting the empirical distribution with wavelet expansions of densities. We use the wavelet frame introduced by [SIAM J. Math. Anal. 38 (2006b) 574-594 (electronic)], the so-called needlets. The expansions are truncated at scale indices no larger than some ${J^{\star}}$, and the $L^p$ distances between those estimates and the null density are computed. One family of tests (called Multiple) is based on the idea of testing the distance from the null for each choice of $J=1,\ldots,{J^{\star}}$, whereas the so-called PlugIn approach is based on the single full ${J^{\star}}$ expansion, but with thresholded wavelet coefficients. We describe the practical implementation of these two procedures and compare them to other methods in the literature. As alternatives to isotropy, we consider both very simple toy models and more realistic nonisotropic models based on Physics-inspired simulations. The Monte Carlo study shows good performance of the Multiple test, even at moderate sample size, for a wide sample of alternative hypotheses and for different choices of the parameter ${J^{\star}}$. On the 69 most energetic events published by the Pierre Auger Collaboration, the needlet-based procedures suggest statistical evidence for anisotropy. Using several values for the parameters of the methods, our procedures yield $p$-values below 1%, but with uncontrolled multiplicity issues. The flexibility of this method and the possibility to modify it to take into account a large variety of extensions of the problem make it an interesting option for future investigation of the origin of ultrahigh energy cosmic rays.

stat.AP

Anisotropic Denoising in Functional Deconvolution Model with Dimension-free Convergence Rates

In the present paper we consider the problem of estimating a periodic $(r+1)$-dimensional function $f$ based on observations from its noisy convolution. We construct a wavelet estimator of $f$, derive minimax lower bounds for the $L^2$-risk when $f$ belongs to a Besov ball of mixed smoothness and demonstrate that the wavelet estimator is adaptive and asymptotically near-optimal within a logarithmic factor, in a wide range of Besov balls. We prove in particular that choosing this type of mixed smoothness leads to rates of convergence which are free of the "curse of dimensionality" and, hence, are higher than usual convergence rates when $r$ is large. The problem studied in the paper is motivated by seismic inversion which can be reduced to solution of noisy two-dimensional convolution equations that allow to draw inference on underground layer structures along the chosen profiles. The common practice in seismology is to recover layer structures separately for each profile and then to combine the derived estimates into a two-dimensional function. By studying the two-dimensional version of the model, we demonstrate that this strategy usually leads to estimators which are less accurate than the ones obtained as two-dimensional functional deconvolutions. Indeed, we show that unless the function $f$ is very smooth in the direction of the profiles, very spatially inhomogeneous along the other direction and the number of profiles is very limited, the functional deconvolution solution has a much better precision compared to a combination of $M$ solutions of separate convolution equations. A limited simulation study in the case of $r=1$ confirms theoretical claims of the paper.

math.ST

Grouping Strategies and Thresholding for High Dimensional Linear Models

The estimation problem in a high regression model with structured sparsity is investigated. An algorithm using a two steps block thresholding procedure called GR-LOL is provided. Convergence rates are produced: they depend on simple coherence-type indices of the Gram matrix -easily checkable on the data- as well as sparsity assumptions of the model parameters measured by a combination of $l_1$ within-blocks with $l_q,q<1$ between-blocks norms. The simplicity of the coherence indicator suggests ways to optimize the rates of convergence when the group structure is not naturally given by the problem and is unknown. In such a case, an auto-driven procedure is provided to determine the regressors groups (number and contents). An intensive practical study compares our grouping methods with the standard LOL algorithm. We prove that the grouping rarely deteriorates the results but can improve them very significantly. GR-LOL is also compared with group-Lasso procedures and exhibits a very encouraging behavior. The results are quite impressive, especially when GR-LOL algorithm is combined with a grouping pre-processing.

math.ST

Thomas Bayes' walk on manifolds

Convergence of the Bayes posterior measure is considered in canonical statistical settings where observations sit on a geometrical object such as a compact manifold, or more generally on a compact metric space verifying some conditions. A natural geometric prior based on randomly rescaled solutions of the heat equation is considered. Upper and lower bound posterior contraction rates are derived.

math.ST

Radon needlet thresholding

We provide a new algorithm for the treatment of the noisy inversion of the Radon transform using an appropriate thresholding technique adapted to a well-chosen new localized basis. We establish minimax results and prove their optimality. In particular, we prove that the procedures provided here are able to attain minimax bounds for any $\mathbb {L}_p$ loss. It s important to notice that most of the minimax bounds obtained here are new to our knowledge. It is also important to emphasize the adaptation properties of our procedures with respect to the regularity (sparsity) of the object to recover and to inhomogeneous smoothness. We perform a numerical study that is of importance since we especially have to discuss the cubature problems and propose an averaging procedure that is mostly in the spirit of the cycle spinning performed for periodic signals.

math.ST

Localized spherical deconvolution

We provide a new algorithm for the treatment of the deconvolution problem on the sphere which combines the traditional SVD inversion with an appropriate thresholding technique in a well chosen new basis. We establish upper bounds for the behavior of our procedure for any $\mathbb {L}_p$ loss. It is important to emphasize the adaptation properties of our procedures with respect to the regularity (sparsity) of the object to recover as well as to inhomogeneous smoothness. We also perform a numerical study which proves that the procedure shows very promising properties in practice as well.

math.ST

A new selection method for high-dimensionial instrumental setting: application to the Growth Rate convergence hypothesis

This paper investigates the problem of selecting variables in regression-type models for an "instrumental" setting. Our study is motivated by empirically verifying the conditional convergence hypothesis used in the economical literature concerning the growth rate. To avoid unnecessary discussion about the choice and the pertinence of instrumental variables, we embed the model in a very high dimensional setting. We propose a selection procedure with no optimization step called LOLA, for Learning Out of Leaders with Adaptation. LOLA is an auto-driven algorithm with two thresholding steps. The consistency of the procedure is proved under sparsity conditions and simulations are conducted to illustrate the practical good performances of LOLA. The behavior of the algorithm is studied when instrumental variables are artificially added without a priori significant connection to the model. Using our algorithm, we provide a solution for modeling the link between the growth rate and the initial level of the gross domestic product and empirically prove the convergence hypothesis.

math.ST

Concentration Inequalities and Confidence Bands for Needlet Density Estimators on Compact Homogeneous Manifolds

Let $X_1,...,X_n$ be a random sample from some unknown probability density $f$ defined on a compact homogeneous manifold $\mathbf M$ of dimension $d \ge 1$. Consider a 'needlet frame' $\{ϕ_{j η}\}$ describing a localised projection onto the space of eigenfunctions of the Laplace operator on $\mathbf M$ with corresponding eigenvalues less than $2^{2j}$, as constructed in \cite{GP10}. We prove non-asymptotic concentration inequalities for the uniform deviations of the linear needlet density estimator $f_n(j)$ obtained from an empirical estimate of the needlet projection $\sum_ηϕ_{j η} \int f ϕ_{j η}$ of $f$. We apply these results to construct risk-adaptive estimators and nonasymptotic confidence bands for the unknown density $f$. The confidence bands are adaptive over classes of differentiable and H\"{older}-continuous functions on $\mathbf M$ that attain their Hölder exponents.

math.ST

Learning Out of Leaders

This paper investigates the estimation problem in a regression-type model. To be able to deal with potential high dimensions, we provide a procedure called LOL, for Learning Out of Leaders with no optimization step. LOL is an auto-driven algorithm with two thresholding steps. A first adaptive thresholding helps to select leaders among the initial regressors in order to obtain a first reduction of dimensionality. Then a second thresholding is performed on the linear regression upon the leaders. The consistency of the procedure is investigated. Exponential bounds are obtained, leading to minimax and adaptive results for a wide class of sparse parameters, with (quasi) no restriction on the number p of possible regressors. An extensive computational experiment is conducted to emphasize the practical good performances of LOL.

math.ST

Inversion of noisy Radon transform by SVD based needlet

A linear method for inverting noisy observations of the Radon transform is developed based on decomposition systems (needlets) with rapidly decaying elements induced by the Radon transform SVD basis. Upper bounds of the risk of the estimator are established in $L^p$ ($1\le p\le \infty$) norms for functions with Besov space smoothness. A practical implementation of the method is given and several examples are discussed.

math.ST

Spin Needlets for Cosmic Microwave Background Polarization Data Analysis

Scalar wavelets have been used extensively in the analysis of Cosmic Microwave Background (CMB) temperature maps. Spin needlets are a new form of (spin) wavelets which were introduced in the mathematical literature by Geller and Marinucci (2008) as a tool for the analysis of spin random fields. Here we adopt the spin needlet approach for the analysis of CMB polarization measurements. The outcome of experiments measuring the polarization of the CMB are maps of the Stokes Q and U parameters which are spin 2 quantities. Here we discuss how to transform these spin 2 maps into spin 2 needlet coefficients and outline briefly how these coefficients can be used in the analysis of CMB polarization data. We review the most important properties of spin needlets, such as localization in pixel and harmonic space and asymptotic uncorrelation. We discuss several statistical applications, including the relation of angular power spectra to the needlet coefficients, testing for non-Gaussianity on polarization data, and reconstruction of the E and B scalar maps.

astro-ph

Needlet algorithms for estimation in inverse problems

We provide a new algorithm for the treatment of inverse problems which combines the traditional SVD inversion with an appropriate thresholding technique in a well chosen new basis. Our goal is to devise an inversion procedure which has the advantages of localization and multiscale analysis of wavelet representations without losing the stability and computability of the SVD decompositions. To this end we utilize the construction of localized frames (termed "needlets") built upon the SVD bases. We consider two different situations: the "wavelet" scenario, where the needlets are assumed to behave similarly to true wavelets, and the "Jacobi-type" scenario, where we assume that the properties of the frame truly depend on the SVD basis at hand (hence on the operator). To illustrate each situation, we apply the estimation algorithm respectively to the deconvolution problem and to the Wicksell problem. In the latter case, where the SVD basis is a Jacobi polynomial basis, we show that our scheme is capable of achieving rates of convergence which are optimal in the $L_2$ case, we obtain interesting rates of convergence for other $L_p$ norms which are new (to the best of our knowledge) in the literature, and we also give a simulation study showing that the NEED-D estimator outperforms other standard algorithms in almost all situations.

math.ST

High Frequency Asymptotics for Wavelet-Based Tests for Gaussianity and Isotropy on the Torus

We prove a CLT for skewness and kurtosis of the wavelets coefficients of a stationary field on the torus. The results are in the framework of the fixed-domain asymptotics, i.e. we refer to observations of a single field which is sampled at higher and higher frequencies. We consider also studentized statistics for the case of an unknown correlation structure. The results are motivated by the analysis of cosmological data or high-frequency financial data sets, with a particular interest towards testing for Gaussianity and isotropy

math.ST