Searcharxiv⌕ Search

arXiv subjects

José R. Berrendero

Publications and source records attributed to José R. Berrendero.

12 recordsLinked to original sources

A Bayesian One-Sample Test for the Mean of Functional Data

In the context of functional data analysis, we propose a Bayesian test to assess whether the mean function of a Gaussian population is identically zero. By leveraging the structure of the reproducing kernel Hilbert space (RKHS) associated with the covariance function of the underlying process, we approximate the mean by a finite linear combination of kernel sections. This finite-dimensional representation arises from the natural assumption that the mean lies in the corresponding RKHS, enabling a parametric Bayesian approach to hypothesis testing through a spike-and-slab prior on the coefficients. Given the absence of a canonical reference measure on infinite-dimensional spaces, the likelihood is expressed via a Radon-Nikodym derivative between induced Gaussian measures. Since the number of kernel sections is unknown, we employ reversible jump Markov chain Monte Carlo (RJMCMC) to explore the trans-dimensional parameter space and estimate the posterior distribution. From the resulting posterior samples we construct estimators of the posterior probability of the null hypothesis and the Bayes factor, providing quantifiable evidence against the null and supporting informed decisions under uncertainty. Furthermore, to validate this methodology we establish near-optimal posterior contraction rates for the proposed test under mild conditions, demonstrate its empirical performance on simulated data sets, and illustrate its application with real climate data.

stat.ME↗

A Bayesian approach to functional regression: theory and computation

We propose a novel Bayesian methodology for inference in functional linear and logistic regression models based on the theory of reproducing kernel Hilbert spaces (RKHS's). We introduce general models that build upon the RKHS generated by the covariance function of the underlying stochastic process, and whose formulation includes as particular cases all finite-dimensional models based on linear combinations of marginals of the process, which can collectively be seen as a dense subspace made of simple approximations. By imposing a suitable prior distribution on this dense functional space we can perform data-driven inference via standard Bayes methodology, estimating the posterior distribution through reversible jump Markov chain Monte Carlo methods. In this context, our contribution is two-fold. First, we derive theoretical results that guarantee strong posterior consistency and contraction at an optimal rate under mild conditions. Second, we show that several prediction strategies stemming from our Bayesian procedure are competitive against other usual alternatives in both simulations and real data sets, including a Bayesian-motivated variable selection method.

stat.ME↗

Relation between PLS and OLS regression in terms of the eigenvalue distribution of the regressor covariance matrix

Partial least squares (PLS) is a dimensionality reduction technique introduced in the field of chemometrics and successfully employed in many other areas. The PLS components are obtained by maximizing the covariance between linear combinations of the regressors and of the target variables. In this work, we focus on its application to scalar regression problems. PLS regression consists in finding the least squares predictor that is a linear combination of a subset of the PLS components. Alternatively, PLS regression can be formulated as a least squares problem restricted to a Krylov subspace. This equivalent formulation is employed to analyze the distance between ${\hat{\boldsymbolβ}\;}_{\mathrm{PLS}}^{\scriptscriptstyle {(L)}}$, the PLS estimator of the vector of coefficients of the linear regression model based on $L$ PLS components, and $\hat{\boldsymbol β}_{\mathrm{OLS}}$, the one obtained by ordinary least squares (OLS), as a function of $L$. Specifically, ${\hat{\boldsymbolβ}\;}_{\mathrm{PLS}}^{\scriptscriptstyle {(L)}}$ is the vector of coefficients in the aforementioned Krylov subspace that is closest to $\hat{\boldsymbol β}_{\mathrm{OLS}}$ in terms of the Mahalanobis distance with respect to the covariance matrix of the OLS estimate. We provide a bound on this distance that depends only on the distribution of the eigenvalues of the regressor covariance matrix. Numerical examples on synthetic and real-world data are used to illustrate how the distance between ${\hat{\boldsymbolβ}\;}_{\mathrm{PLS}}^{\scriptscriptstyle {(L)}}$ and $\hat{\boldsymbol β}_{\mathrm{OLS}}$ depends on the number of clusters in which the eigenvalues of the regressor covariance matrix are grouped.

stat.ME↗

On functional logistic regression: some conceptual issues

The main ideas behind the classical multivariate logistic regression model make sense when translated to the functional setting, where the explanatory variable $X$ is a function and the response $Y$ is binary. However, some important technical issues appear (or are aggravated with respect to those of the multivariate case) due to the functional nature of the explanatory variable. First, the mere definition of the model can be questioned: while most approaches so far proposed rely on the $L_2$-based model, we suggest an alternative (in some sense, more general) approach, based on the theory of Reproducing Kernel Hilbert Spaces (RKHS). The validity conditions of such RKHS-based model, as well as its relation with the $L_2$-based one are investigated and made explicit in two formal results. Some relevant particular cases are considered as well. Second we show that, under very general conditions, the maximum likelihood (ML) of the logistic model parameters fail to exist in the functional case. Third, on a more positive side, we suggest an RKHS-based restricted version of the ML estimator. This is a methodological paper, aimed at a better understanding of the functional logistic model, rather than focusing on numerical and practical issues.

math.ST↗

On a general definition of the functional linear model

A general formulation of the linear model with functional (random) explanatory variable $X = X(t), t \in T$ , and scalar response Y is proposed. It includes the standard functional linear model, based on the inner product in the space $L^2[0,1]$, as a particular case. It also includes all models in which Y is assumed to be (up to an additive noise) a linear combination of a finite or countable collections of marginal variables X(t_j), with $t_j\in T$ or a linear combination of a finite number of linear projections of X. This general formulation can be interpreted in terms of the RKHS space generated by the covariance function of the process X(t). Some consistency results are proved. A few experimental results are given in order to show the practical interest of considering, in a unified framework, linear models based on a finite number of marginals $X(t_j)$ of the process $X(t)$.

math.ST↗

On Mahalanobis distance in functional settings

Mahalanobis distance is a classical tool in multivariate analysis. We suggest here an extension of this concept to the case of functional data. More precisely, the proposed definition concerns those statistical problems where the sample data are real functions defined on a compact interval of the real line. The obvious difficulty for such a functional extension is the non-invertibility of the covariance operator in infinite-dimensional cases. Unlike other recent proposals, our definition is suggested and motivated in terms of the Reproducing Kernel Hilbert Space (RKHS) associated with the stochastic process that generates the data. The proposed distance is a true metric; it depends on a unique real smoothing parameter which is fully motivated in RKHS terms. Moreover, it shares some properties of its finite dimensional counterpart: it is invariant under isometries, it can be consistently estimated from the data and its sampling distribution is known under Gaussian models. An empirical study for two statistical applications, outliers detection and binary classification, is included. The obtained results are quite competitive when compared to other recent proposals of the literature.

stat.ME↗

An RKHS model for variable selection in functional regression

A mathematical model for variable selection in functional regression models with scalar response is proposed. By "variable selection" we mean a procedure to replace the whole trajectories of the functional explanatory variables with their values at a finite number of carefully selected instants (or "impact points"). The basic idea of our approach is to use the Reproducing Kernel Hilbert Space (RKHS) associated with the underlying process, instead of the more usual L2[0,1] space, in the definition of the linear model. This turns out to be especially suitable for variable selection purposes, since the finite-dimensional linear model based on the selected "impact points" can be seen as a particular case of the RKHS-based linear functional model. In this framework, we address the consistent estimation of the optimal design of impact points and we check, via simulations and real data examples, the performance of the proposed method.

stat.ME↗

On the use of reproducing kernel Hilbert spaces in functional classification

The Hájek-Feldman dichotomy establishes that two Gaussian measures are either mutually absolutely continuous with respect to each other (and hence there is a Radon-Nikodym density for each measure with respect to the other one) or mutually singular. Unlike the case of finite dimensional Gaussian measures, there are non-trivial examples of both situations when dealing with Gaussian stochastic processes. This paper provides: (a) Explicit expressions for the optimal (Bayes) rule and the minimal classification error probability in several relevant problems of supervised binary classification of mutually absolutely continuous Gaussian processes. The approach relies on some classical results in the theory of Reproducing Kernel Hilbert Spaces (RKHS). (b) An interpretation, in terms of mutual singularity, for the "near perfect classification" phenomenon described by Delaigle and Hall (2012). We show that the asymptotically optimal rule proposed by these authors can be identified with the sequence of optimal rules for an approximating sequence of classification problems in the absolutely continuous case. (c) A new model-based method for variable selection in binary classification problems, which arises in a very natural way from the explicit knowledge of the RN-derivatives and the underlying RKHS structure. Different classifiers might be used from the selected variables. In particular, the classical, linear finite-dimensional Fisher rule turns out to be consistent under some standard conditions on the underlying functional model.

stat.ME↗

The mRMR variable selection method: a comparative study for functional data

The use of variable selection methods is particularly appealing in statistical problems with functional data. The obvious general criterion for variable selection is to choose the `most representative' or `most relevant' variables. However, it is also clear that a purely relevance-oriented criterion could lead to select many redundant variables. The mRMR (minimum Redundance Maximum Relevance) procedure, proposed by Ding and Peng (2005) and Peng et al. (2005) is an algorithm to systematically perform variable selection, achieving a reasonable trade-off between relevance and redundancy. In its original form, this procedure is based on the use of the so-called mutual information criterion to assess relevance and redundancy. Keeping the focus on functional data problems, we propose here a modified version of the mRMR method, obtained by replacing the mutual information by the new association measure (called distance correlation) suggested by Székely et al. (2007). We have also performed an extensive simulation study, including 1600 functional experiments (100 functional models $\times$ 4 sample sizes $\times$ 4 classifiers) and three real-data examples aimed at comparing the different versions of the mRMR methodology. The results are quite conclusive in favor of the new proposed alternative.

stat.ME↗

Variable selection in functional data classification: a maxima-hunting proposal

Variable selection is considered in the setting of supervised binary classification with functional data $\{X(t),\ t\in[0,1]\}$. By "variable selection" we mean any dimension-reduction method which leads to replace the whole trajectory $\{X(t),\ t\in[0,1]\}$, with a low-dimensional vector $(X(t_1),\ldots,X(t_k))$ still keeping a similar classification error. Our proposal for variable selection is based on the idea of selecting the local maxima $(t_1,\ldots,t_k)$ of the function ${\mathcal V}_X^2(t)={\mathcal V}^2(X(t),Y)$, where ${\mathcal V}$ denotes the "distance covariance" association measure for random variables due to Székely, Rizzo and Bakirov (2007). This method provides a simple natural way to deal with the relevance vs. redundancy trade-off which typically appears in variable selection. This paper includes (a) Some theoretical motivation: a result of consistent estimation on the maxima of ${\mathcal V}_X^2$ is shown. We also show different theoretical models for the underlying process $X(t)$ under which the relevant information in concentrated in the maxima of ${\mathcal V}_X^2$. (b) An extensive empirical study, including about 400 simulated models and real data examples, aimed at comparing our variable selection method with other standard proposals for dimension reduction.

stat.ME↗

A geometrically motivated parametric model in manifold estimation,

The general aim of manifold estimation is reconstructing, by statistical methods, an $m$-dimensional compact manifold $S$ on ${\mathbb R}^d$ (with $m\leq d$) or estimating some relevant quantities related to the geometric properties of $S$. We will assume that the sample data are given by the distances to the $(d-1)$-dimensional manifold $S$ from points randomly chosen on a band surrounding $S$, with $d=2$ and $d=3$. The point in this paper is to show that, if $S$ belongs to a wide class of compact sets (which we call \it sets with polynomial volume\rm), the proposed statistical model leads to a relatively simple parametric formulation. In this setup, standard methodologies (method of moments, maximum likelihood) can be used to estimate some interesting geometric parameters, including curvatures and Euler characteristic. We will particularly focus on the estimation of the $(d-1)$-dimensional boundary measure (in Minkowski's sense) of $S$. It turns out, however, that the estimation problem is not straightforward since the standard estimators show a remarkably pathological behavior: while they are consistent and asymptotically normal, their expectations are infinite. The theoretical and practical consequences of this fact are discussed in some detail.

math.ST↗

On the maximum bias functions of MM-estimates and constrained M-estimates of regression

We derive the maximum bias functions of the MM-estimates and the constrained M-estimates or CM-estimates of regression and compare them to the maximum bias functions of the S-estimates and the $τ$-estimates of regression. In these comparisons, the CM-estimates tend to exhibit the most favorable bias-robustness properties. Also, under the Gaussian model, it is shown how one can construct a CM-estimate which has a smaller maximum bias function than a given S-estimate, that is, the resulting CM-estimate dominates the S-estimate in terms of maxbias and, at the same time, is considerably more efficient.

math.ST↗