Gasper: GrAph Signal ProcEssing in R
We present a short tutorial on to the use of the R gasper package. Gasper is a package dedicated to signal processing on graphs. It also provides an interface to the SuiteSparse Matrix Collection.
arXiv subjects
Publications and source records attributed to Fabien Navarro.
We present a short tutorial on to the use of the R gasper package. Gasper is a package dedicated to signal processing on graphs. It also provides an interface to the SuiteSparse Matrix Collection.
The COVID-19 pandemic has taken the world by storm with its high infection rate. Investigating its geographical disparities has paramount interest in order to gauge its relationships with political decisions, economic indicators, or mental health. This paper focuses on clustering the daily death rates reported in several regions of Europe and the United States over eight months. Several methods have been developed to cluster such functional data. However, these methods are not translation-invariant and thus cannot handle different times of arrivals of the disease, nor can they consider external covariates and so are unable to adjust for the population risk factors of each region. We propose a novel three-step clustering method to circumvent these issues. As a first step, feature extraction is performed by translation-invariant wavelet decomposition which permits to deal with the different onsets. As a second step, single-index regression is used to neutralize disparities caused by population risk factors. As a third step, a nonparametric mixture is fitted on the regression residuals to achieve the region clustering. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available online.
Over the last decade, signal processing on graphs has become a very active area of research. Specifically, the number of applications, for instance in statistical or deep learning, using frames built from graphs, such as wavelets on graphs, has increased significantly. We consider in particular the case of signal denoising on graphs via a data-driven wavelet tight frame methodology. This adaptive approach is based on a threshold calibrated using Stein's unbiased risk estimate adapted to a tight-frame representation. We make it scalable to large graphs using Chebyshev-Jackson polynomial approximations, which allow fast computation of the wavelet coefficients, without the need to compute the Laplacian eigendecomposition. However, the overcomplete nature of the tight-frame, transforms a white noise into a correlated one. As a result, the covariance of the transformed noise appears in the divergence term of the SURE, thus requiring the computation and storage of the frame, which leads to an impractical calculation for large graphs. To estimate such covariance, we develop and analyze a Monte-Carlo strategy, based on the fast transformation of zero mean and unit variance random variables. This new data-driven denoising methodology finds a natural application in differential privacy. A comprehensive performance analysis is carried out on graphs of varying size, from real and simulated data.
We propose a new point of view in the study of Fourier analysis on graphs, taking advantage of localization in the Fourier domain. For a signal $f$ on vertices of a weighted graph $\mathcal{G}$ with Laplacian matrix $\mathcal{L}$, standard Fourier analysis of $f$ relies on the study of functions $g(\mathcal{L})f$ for some filters $g$ on $I_\mathcal{L}$, the smallest interval containing the Laplacian spectrum ${\mathrm sp}(\mathcal{L}) \subset I_\mathcal{L}$. We show that for carefully chosen partitions $I_\mathcal{L} = \sqcup_{1\leq k\leq K} I_k$ ($I_k \subset I_\mathcal{L}$), there are many advantages in understanding the collection $(g(\mathcal{L}_{I_k})f)_{1\leq k\leq K}$ instead of $g(\mathcal{L})f$ directly, where $\mathcal{L}_I$ is the projected matrix $P_I(\mathcal{L})\mathcal{L}$. First, the partition provides a convenient modelling for the study of theoretical properties of Fourier analysis and allows for new results in graph signal analysis (\emph{e.g.} noise level estimation, Fourier support approximation). We extend the study of spectral graph wavelets to wavelets localized in the Fourier domain, called LocLets, and we show that well-known frames can be written in terms of LocLets. From a practical perspective, we highlight the interest of the proposed localized Fourier analysis through many experiments that show significant improvements in two different tasks on large graphs, noise level estimation and signal denoising. Moreover, efficient strategies permit to compute sequence $(g(\mathcal{L}_{I_k})f)_{1\leq k\leq K}$ with the same time complexity as for the computation of $g(\mathcal{L})f$.
This paper is devoted to adaptive signal denoising in the context of Graph Signal Processing (GSP) using Spectral Graph Wavelet Transform (SGWT). This issue is addressed \emph{via} a data-driven thresholding process in the transformed domain by optimizing the parameters in the sense of the Mean Square Error (MSE) using the Stein's Unbiased Risk Estimator (SURE). The SGWT considered is built upon a partition of unity making the transform semi-orthogonal so that the optimization can be performed in the transformed domain. However, since the SGWT is over-complete, the divergence term in the SURE needs to be computed in the context of correlated noise. Two thresholding strategies called coordinatewise and block thresholding process are investigated. For each of them, the SURE is derived for a whole family of elementary thresholding functions among which the soft threshold and the James-Stein threshold. This multi-scales analysis shows better performance than the most recent methods from the literature. That is illustrated numerically for a series of signals on different graphs.
In this paper, we consider an unknown functional estimation problem in a general nonparametric regression model with the feature of having both multiplicative and additive noise.We propose two new wavelet estimators in this general context. We prove that they achieve fast convergence rates under the mean integrated square error over Besov spaces. The obtained rates have the particularity of being established under weak conditions on the model. A numerical study in a context comparable to stochastic frontier estimation (with the difference that the boundary is not necessarily a production function) supports the theory.
Motivated by the analysis of accelerometer data, we introduce a specific finite mixture of hidden Markov models with particular characteristics that adapt well to the specific nature of this type of data. Our model allows for the computation of statistics that characterize the physical activity of a subject (\emph{e.g.}, the mean time spent at different activity levels and the probability of the transition between two activity levels) without specifying the activity levels in advance but by estimating them from the data. In addition, this approach allows the heterogeneity of the population to be taken into account and subpopulations with homogeneous physical activity behavior to be defined. We prove that, under mild assumptions, this model implies that the probability of misclassifying a subject decreases at an exponential decay with the length of its measurement sequence. Model identifiability is also investigated. We also report a comprehensive suite of numerical simulations to support our theoretical findings. Method is motivated by and applied to the PAT study.
We emphasize that it is possible to improve the principle of unbiased risk estimation for model selection by addressing excess risk deviations in the design of penalization procedures. Indeed, we propose a modification of Akaike's Information Criterion that avoids overfitting, even when the sample size is small. We call this correction an over-penalization procedure. As proof of concept, we show the nonasymptotic optimality of our histogram selection procedure in density estimation by establishing sharp oracle inequalities for the Kullback-Leibler divergence. One of the main features of our theoretical results is that they include the estimation of unbounded logdensities. To do so, we prove several analytical and probabilistic lemmas that are of independent interest. In an experimental study, we also demonstrate state-of-the-art performance of our over-penalization criterion for bin size selection, in particular outperforming AICc procedure.
We investigate the optimality for model selection of the so-called slope heuristics, $V$-fold cross-validation and $V$-fold penalization in a heteroscedastic with random design regression context. We consider a new class of linear models that we call strongly localized bases and that generalize histograms, piecewise polynomials and compactly supported wavelets. We derive sharp oracle inequalities that prove the asymptotic optimality of the slope heuristics---when the optimal penalty shape is known---and $V$ -fold penalization. Furthermore, $V$-fold cross-validation seems to be suboptimal for a fixed value of $V$ since it recovers asymptotically the oracle learned from a sample size equal to $1-V^{-1}$ of the original amount of data. Our results are based on genuine concentration inequalities for the true and empirical excess risks that are of independent interest. We show in our experiments the good behavior of the slope heuristics for the selection of linear wavelet models. Furthermore, $V$-fold cross-validation and $V$-fold penalization have comparable efficiency.
We observe $n$ heteroscedastic stochastic processes $\{Y_v(t)\}_{v}$, where for any $v\in\{1,\ldots,n\}$ and $t \in [0,1]$, $Y_v(t)$ is the convolution product of an unknown function $f$ and a known blurring function $g_v$ corrupted by Gaussian noise. Under an ordinary smoothness assumption on $g_1,\ldots,g_n$, our goal is to estimate the $d$-th derivatives (in weak sense) of $f$ from the observations. We propose an adaptive estimator based on wavelet block thresholding, namely the "BlockJS estimator". Taking the mean integrated squared error (MISE), our main theoretical result investigates the minimax rates over Besov smoothness spaces, and shows that our block estimator can achieve the optimal minimax rate, or is at least nearly-minimax in the least favorable situation. We also report a comprehensive suite of numerical simulations to support our theoretical findings. The practical performance of our block estimator compares very favorably to existing methods of the literature on a large set of test functions.
We investigate the estimation of a weighted density taking the form $g=w(F)f$, where $f$ denotes an unknown density, $F$ the associated distribution function and $w$ is a known (non-negative) weight. Such a class encompasses many examples, including those arising in order statistics or when $g$ is related to the maximum or the minimum of $N$ (random or fixed) independent and identically distributed (\iid) random variables. We here construct a new adaptive non-parametric estimator for $g$ based on a plug-in approach and the wavelets methodology. For a wide class of models, we prove that it attains fast rates of convergence under the $\mathbb{L}_p$ risk with $p\ge 1$ (not only for $p = 2$ corresponding to the mean integrated squared error) over Besov balls. The theoretical findings are illustrated through several simulations.