SearcharxivSearch

arXiv subjects

Pamela Llop

Publications and source records attributed to Pamela Llop.

9 recordsLinked to original sources

A semiparametric autorregresive spatial prediction model

In this paper we propose a semiparametric spatial autoregressive model that combines a linear covariate component with a nonparametrically estimated spatial term, allowing flexible dependence modeling without restrictive covariance structure while preserving interpretability. We establish asymptotic properties, including consistency and asymptotic normality, and evaluate performance through simulations and real data. Results show competitive predictive accuracy relative to geostatistical methods and improved interpretability compared to spatial econometric models.

stat.ME

Value-Aware Product Recommendation by Customer Segmentation using a suitable High-Dimensional Similarity Measure

This paper presents a novel value-aware approach to product recommendation that simultaneously addresses the high dimensionality and sparsity of user-item data while explicitly incorporating the contribution of each product and user to overall sales revenue. The proposed framework encodes revenue contributions in the user-item matrix and computes customer similarity directly on this basis using suitable distance measures. This enables the segmentation of users according to the revenue-based similarity of their purchase baskets and supports recommendations aligned with profitability objectives. We compare conventional similarity metrics with a novel alternative tailored to high-dimensional contexts and propose three recommendation strategies based on revenue share, product popularity, and expected profit generation. The effectiveness of the proposed method is validated through simulation experiments and a real-world application using the UCI Online Retail dataset.

cs.IR

Sufficient dimension reduction for regression with spatially correlated errors: application to prediction

In this paper, we address the problem of predicting a response variable in the context of both, spatially correlated and high-dimensional data. To reduce the dimensionality of the predictor variables, we apply the sufficient dimension reduction (SDR) paradigm, which reduces the predictor space while retaining relevant information about the response. To achieve this, we impose two different spatial models on the inverse regression: the separable spatial covariance model (SSCM) and the spatial autoregressive error model (SEM). For these models, we derive maximum likelihood estimators for the reduction and use them to predict the response via nonparametric rules for forward regression. Through simulations and real data applications, we demonstrate the effectiveness of our approach for spatial data prediction.

stat.ME

Sufficient reductions in regression with mixed predictors

Most data sets comprise of measurements on continuous and categorical variables. In regression and classification Statistics literature, modeling high-dimensional mixed predictors has received limited attention. In this paper we study the general regression problem of inferring on a variable of interest based on high dimensional mixed continuous and binary predictors. The aim is to find a lower dimensional function of the mixed predictor vector that contains all the modeling information in the mixed predictors for the response, which can be either continuous or categorical. The approach we propose identifies sufficient reductions by reversing the regression and modeling the mixed predictors conditional on the response. We derive the maximum likelihood estimator of the sufficient reductions, asymptotic tests for dimension, and a regularized estimator, which simultaneously achieves variable (feature) selection and dimension reduction (feature extraction). We study the performance of the proposed method and compare it with other approaches through simulations and real data examples.

math.ST

Comparing statistical methods to predict leptospirosis incidence using hydro-climatic covariables

Leptospiroris, the infectious disease caused by the spirochete bacteria Leptospira interrogans, constitutes an important public health problem all over the world. In Argentina, some regions present climate and geographic characteristics that favors the habitat of the bacteria Leptospira, whose survival strongly depends on climatic factors. For this reason, regional public health systems should include, as a main factor, the incidence of the disease in order to improve the prediction of potential outbreaks, helping to stop or delay the virus transmission. The classic methods used to perform this kind of predictions are based in autoregressive time series tools which, as it is well known, perform poorly when the data do not meet their requirements. Recently, several nonparametric methods have been introduced to deal with those problems. In this work, we compare a semiparametric method, called Semi-Functional Partial Linear Regression (SFPLR) with the classic ARIMA and a new alternative ARIMAX, in order to select the best predictive tool for the incidence of leptospirosis in the Argentinian Litoral region. In particular, SFPLR and ARIMAX are methods that allow the use of (hydrometeorological) covariables which could improve the prediction of outbreaks of leptospirosis.

stat.AP

Supervised dimension reduction for ordinal predictors

In applications involving ordinal predictors, common approaches to reduce dimensionality are either extensions of unsupervised techniques such as principal component analysis, or variable selection procedures that rely on modeling the regression function. In this paper, a supervised dimension reduction method tailored to ordered categorical predictors is introduced. It uses a model-based dimension reduction approach, inspired by extending sufficient dimension reductions to the context of latent Gaussian variables. The reduction is chosen without modeling the response as a function of the predictors and does not impose any distributional assumption on the response or on the response given the predictors. A likelihood-based estimator of the reduction is derived and an iterative expectation-maximization type algorithm is proposed to alleviate the computational load and thus make the method more practical. A regularized estimator, which simultaneously achieves variable selection and dimension reduction, is also presented. Performance of the proposed method is evaluated through simulations and a real data example for socioeconomic index construction, comparing favorably to widespread use techniques.

math.ST

On the classification problem for Poisson Point Processes

We study the binary classification problem for Poisson point processes, which are allowed to take values in a general metric space. The problem is tackled in two different ways: estimating nonparametricaly the intensity functions of the processes (and then plugged into a deterministic formula which expresses the regression function in terms of the intensities), and performing the classical $k$ nearest neighbor rule by introducing a suitable distance between patterns of points. In the first approach we prove the consistency of the estimated intensity so that the rule turns out to be also consistent. For the $k$-NN classifier, we prove that the regression function fulfils the so called "Besicovitch condition", usually required for the consistency of the classical classification rules. The theoretical findings are illustrated on simulated data, where in one case the $k$-NN rule outperforms the first approach.

math.ST

A nonlinear aggregation type classifier

We introduce a nonlinear aggregation type classifier for functional data defined on a separable and complete metric space. The new rule is built up from a collection of $M$ arbitrary training classifiers. If the classifiers are consistent, then so is the aggregation rule. Moreover, asymptotically the aggregation rule behaves as well as the best of the $M$ classifiers. The results of a small simulation are reported both, for high dimensional and functional data, and a real data example is analyzed.

math.ST

An optimal aggregation type classifier

We introduce a nonlinear aggregation type classifier for functional data defined on a separable and complete metric space. The new rule is built up from a collection of $M$ arbitrary training classifiers. If the classifiers are consistent, then so is the aggregation rule. Moreover, asymptotically the aggregation rule behaves as well as the best of the $M$ classifiers. The results of a small si\-mu\-lation are reported both, for high dimensional and functional data.

math.ST