SearcharxivSearch

arXiv subjects

Liliana Forzani

Publications and source records attributed to Liliana Forzani.

14 recordsLinked to original sources

On Tent Spaces for the Gaussian Measure

Following the scheme of tent spaces in classical harmonic analysis developed by R. Coifman, Y. Meyer, and E. Stein in \cite{cms}, we succeed in doing so for the Gaussian setting. In \cite{MNP}, part of this theory (an atomic decomposition) is developed for a specific tent space where functions are defined just in a proper subset of $\mathbb{R}^{n+1}_+,$ and without the use of an area function. In the present paper, using a variation of the area function considered in \cite{FSU}, we define the Gaussian area function and Gaussian tent spaces and prove both their atomic decompositions and the characterization of their dual spaces. Some applications are also considered.

math.AP

Sufficient dimension reduction for regression with spatially correlated errors: application to prediction

In this paper, we address the problem of predicting a response variable in the context of both, spatially correlated and high-dimensional data. To reduce the dimensionality of the predictor variables, we apply the sufficient dimension reduction (SDR) paradigm, which reduces the predictor space while retaining relevant information about the response. To achieve this, we impose two different spatial models on the inverse regression: the separable spatial covariance model (SSCM) and the spatial autoregressive error model (SEM). For these models, we derive maximum likelihood estimators for the reduction and use them to predict the response via nonparametric rules for forward regression. Through simulations and real data applications, we demonstrate the effectiveness of our approach for spatial data prediction.

stat.ME

Asymptotic results for nonparametric regression estimators after sufficient dimension reduction estimation

Prediction, in regression and classification, is one of the main aims in modern data science. When the number of predictors is large, a common first step is to reduce the dimension of the data. Sufficient dimension reduction (SDR) is a well established paradigm of reduction that keeps all the relevant information in the covariates X that is necessary for the prediction of Y . In practice, SDR has been successfully used as an exploratory tool for modelling after estimation of the sufficient reduction. Nevertheless, even if the estimated reduction is a consistent estimator of the population, there is no theory that supports this step when non-parametric regression is used in the imputed estimator. In this paper, we show that the asymptotic distribution of the non-parametric regression estimator is the same regardless if the true SDR or its estimator is used. This result allows making inferences, for example, computing confidence intervals for the regression function avoiding the curse of dimensionality.

stat.ME

Sufficient reductions in regression with mixed predictors

Most data sets comprise of measurements on continuous and categorical variables. In regression and classification Statistics literature, modeling high-dimensional mixed predictors has received limited attention. In this paper we study the general regression problem of inferring on a variable of interest based on high dimensional mixed continuous and binary predictors. The aim is to find a lower dimensional function of the mixed predictor vector that contains all the modeling information in the mixed predictors for the response, which can be either continuous or categorical. The approach we propose identifies sufficient reductions by reversing the regression and modeling the mixed predictors conditional on the response. We derive the maximum likelihood estimator of the sufficient reductions, asymptotic tests for dimension, and a regularized estimator, which simultaneously achieves variable (feature) selection and dimension reduction (feature extraction). We study the performance of the proposed method and compare it with other approaches through simulations and real data examples.

math.ST

Envelopes for multivariate linear regression with linearly constrained coefficients

A constrained multivariate linear model is a multivariate linear model with the columns of its coefficient matrix constrained to lie in a known subspace. This class of models includes those typically used to study growth curves and longitudinal data. Envelope methods have been proposed to improve estimation efficiency in the class of unconstrained multivariate linear models, but have not yet been developed for constrained models that we develop in this article. We first compare the standard envelope estimator based on an unconstrained multivariate model with the standard estimator arising from a constrained multivariate model in terms of bias and efficiency. Then, to further improve efficiency, we propose a novel envelope estimator based on a constrained multivariate model. Novel envelope-based testing methods are also proposed. We provide support for our proposals by simulations and by studying the classical dental data and data from the China Health and Nutrition Survey and a study of probiotic capacity to reduced Salmonella infection.

stat.ME

Fundamentals of path analysis in the social sciences

Motivated by a recent series of diametrically opposed articles on the relative value of statistical methods for the analysis of path diagrams in the social sciences, we discuss from a primarily theoretical perspective selected fundamental aspects of path modeling and analysis based on a common re reflexive setting. Since there is a paucity of technical support evident in the debate, our aim is to connect it to mainline statistics literature and to address selected foundational issues that may help move the discourse. We do not intend to advocate for or against a particular method or analysis philosophy.

stat.ME

Supervised dimension reduction for ordinal predictors

In applications involving ordinal predictors, common approaches to reduce dimensionality are either extensions of unsupervised techniques such as principal component analysis, or variable selection procedures that rely on modeling the regression function. In this paper, a supervised dimension reduction method tailored to ordered categorical predictors is introduced. It uses a model-based dimension reduction approach, inspired by extending sufficient dimension reductions to the context of latent Gaussian variables. The reduction is chosen without modeling the response as a function of the predictors and does not impose any distributional assumption on the response or on the response given the predictors. A likelihood-based estimator of the reduction is derived and an iterative expectation-maximization type algorithm is proposed to alleviate the computational load and thus make the method more practical. A regularized estimator, which simultaneously achieves variable selection and dimension reduction, is also presented. Performance of the proposed method is evaluated through simulations and a real data example for socioeconomic index construction, comparing favorably to widespread use techniques.

math.ST

Asymptotic theory for maximum likelihood estimates in reduced-rank multivariate generalised linear models

Reduced-rank regression is a dimensionality reduction method with many applications. The asymptotic theory for reduced rank estimators of parameter matrices in multivariate linear models has been studied extensively. In contrast, few theoretical results are available for reduced-rank multivariate generalised linear models. We develop M-estimation theory for concave criterion functions that are maximised over parameters spaces that are neither convex nor closed. These results are used to derive the consistency and asymptotic distribution of maximum likelihood estimators in reduced-rank multivariate generalised linear models, when the response and predictor vectors have a joint distribution. We illustrate our results in a real data classification problem with binary covariates.

math.ST

On the classification problem for Poisson Point Processes

We study the binary classification problem for Poisson point processes, which are allowed to take values in a general metric space. The problem is tackled in two different ways: estimating nonparametricaly the intensity functions of the processes (and then plugged into a deterministic formula which expresses the regression function in terms of the intensities), and performing the classical $k$ nearest neighbor rule by introducing a suitable distance between patterns of points. In the first approach we prove the consistency of the estimated intensity so that the rule turns out to be also consistent. For the $k$-NN classifier, we prove that the regression function fulfils the so called "Besicovitch condition", usually required for the consistency of the classical classification rules. The theoretical findings are illustrated on simulated data, where in one case the $k$-NN rule outperforms the first approach.

math.ST

Algorithms for Envelope Estimation II

We propose a new algorithm for envelope estimation, along with a new root n consistent method for computing starting values. The new algorithm, which does not require optimization over a Grassmannian, is shown by simulation to be much faster and typically more accurate that the best existing algorithm proposed by Cook and Zhang (2015c).

stat.ME

Estimating sufficient reductions of the predictors in abundant high-dimensional regressions

We study the asymptotic behavior of a class of methods for sufficient dimension reduction in high-dimension regressions, as the sample size and number of predictors grow in various alignments. It is demonstrated that these methods are consistent in a variety of settings, particularly in abundant regressions where most predictors contribute some information on the response, and oracle rates are possible. Simulation results are presented to support the theoretical conclusion.

math.ST

On convex regression estimators

A new nonparametric estimator of a convex regression function in any dimension is proposed and its convergence properties are studied. We start by using any estimator of the regression function and we \emph{convexify} it by taking the convex envelope of a sample of the approximation obtained. We prove that the uniform rate of convergence of the estimator is maintained after the convexification is applied. The finite sample properties of the new estimator are investigated by means of a simulation study and the application of the new method is demonstrated in examples.

math.ST

Principal Fitted Components for Dimension Reduction in Regression

We provide a remedy for two concerns that have dogged the use of principal components in regression: (i) principal components are computed from the predictors alone and do not make apparent use of the response, and (ii) principal components are not invariant or equivariant under full rank linear transformation of the predictors. The development begins with principal fitted components [Cook, R. D. (2007). Fisher lecture: Dimension reduction in regression (with discussion). Statist. Sci. 22 1--26] and uses normal models for the inverse regression of the predictors on the response to gain reductive information for the forward regression of interest. This approach includes methodology for testing hypotheses about the number of components and about conditional independencies among the predictors.

stat.ME

On the maximal function for the generalized Ornstein-Uhlenbeck semigroup

In this note we consider the maximal function for the generalized Ornstein-Uhlenbeck semigroup in $\RR$ associated with the generalized Hermite polynomials $\{H_n^μ\}$ and prove that it is weak type (1,1) with respect to $dλ_μ(x) = |x|^{2μ}e^{-|x|^2} dx,$ for $μ>-1/2$ as well as bounded on $L^p(dλ_μ) $ for $p>1$

math.CA