Searcharxiv⌕ Search

arXiv subjects

Marcos Oliveira Prates

Publications and source records attributed to Marcos Oliveira Prates.

9 recordsLinked to original sources

Extensions in Semiparametric Geostatistical Models

In spatial statistics, the incorrect selection of an appropriate covariance function may lead to inference errors and confidence underestimation. Motivated by such restrictions, we introduce and evaluate a flexible semiparametric approach for estimating spatial covariance functions based on Bernstein polynomials. The proposed formulation is general and applicable to classes of models that incorporate latent spatial effects in georeferenced data, such as Spatial Generalized Linear Mixed Models. Empirical validation was conducted via Monte Carlo simulations, and model fitting was performed using Bayesian inference via Markov chain Monte Carlo. Simulated scenarios demonstrated the model's ability to recover structural covariance configurations with low bias and high parameter precision. The practical applicability of the methodology was tested using real abundance data for American Robin (Turdus migratorius) from the North American Breeding Bird Survey. The proposed model, featuring a Negative Binomial structure, yielded satisfactory results, efficiently capturing the overdispersion inherent in the count data. The estimated range parameter of 381.48 km revealed that the species' spatial dependence operates at a regional scale, suggesting that unobserved ecological processes act homogeneously within this radius of environmental influence. Additionally, predictive validation using an independent sample (n_pred = 34) demonstrated the model's strong generalization capability via Bayesian Kriging, producing point projections that closely matched observed values and well-calibrated prediction intervals. It is concluded that the proposed approach represents a robust methodological advancement, establishing itself as a flexible and efficient tool.

stat.ME↗

Spatial Confounding: A review of concepts, challenges, and current approaches

Spatial confounding is a persistent challenge in spatial statistics, influencing the validity of statistical inference in models that analyze spatially-structured data. The concept has been interpreted in various ways but is broadly defined as bias in estimates arising from unmeasured spatial variation. In this paper we review definitions, classical spatial models, and recent methodological advances, including approaches from spatial statistics and causal inference. We provide an unified view of the many available approaches for areal as well as geostatistical data and discuss their relative merits both theoretically and empirically with a head-to-head comparison on real datasets. Finally, we leverage the results of the empirical comparisons to discuss directions for future research.

stat.ME↗

On the Validity of Isotropic Covariance Functions for Set-indexed Random Fields

Distances between sets arise naturally when modeling stochastic dependence on collections of spatial supports, including settings with point-referenced and areal observations. However, commonly used constructions of distances on sets, including those derived from the Hausdorff distance, generally fail to be conditionally negative definite, precluding their use in isotropic covariance models. We propose the ball--Hausdorff distance, defined as the Hausdorff distance between the minimum enclosing balls of bounded sets in a metric space. For length spaces, we derive an explicit representation of this distance in terms of the associated centers and radii. We show that the ball--Hausdorff distance is conditionally negative definite whenever the underlying metric is conditionally negative definite. By Schoenberg's theorem, this implies an isometric embedding into a Hilbert space and guarantees the validity of broad classes of isotropic covariance functions, including the Matérn and powered exponential families, for set-indexed random fields. The construction reduces dependence between sets to low-dimensional geometric summaries, leading to substantial simplifications in covariance evaluation.

stat.ME↗

Statistical Inferences and Predictions for Areal Data and Spatial Data Fusion with Hausdorff--Gaussian Processes

Accurate modeling of spatial dependence is pivotal in analyzing spatial data, influencing parameter estimation and predictions. The spatial structure of the data significantly impacts valid statistical inference. Existing models for areal data often rely on adjacency matrices, struggling to differentiate between polygons of varying sizes and shapes. Conversely, data fusion models rely on computationally intensive numerical integrals, presenting challenges for moderately large datasets. In response to these issues, we propose the Hausdorff-Gaussian process (HGP), a versatile model utilizing the Hausdorff distance to capture spatial dependence in both point and areal data. Integration into generalized linear mixed-effects models enhances its applicability, particularly in addressing data fusion challenges. We validate our approach through a comprehensive simulation study and application to two real-world scenarios: one involving areal data and another demonstrating its effectiveness in data fusion. The results suggest that the HGP is competitive with specialized models regarding goodness-of-fit and prediction performances. In summary, the HGP offers a flexible and robust solution for modeling spatial data of various types and shapes, with potential applications spanning fields such as public health and climate science.

stat.ME↗

Imputation of missing data using multivariate Gaussian Linear Cluster-Weighted Modeling

Missing data arises when certain values are not recorded or observed for variables of interest. However, most of the statistical theory assume complete data availability. To address incomplete databases, one approach is to fill the gaps corresponding to the missing information based on specific criteria, known as imputation. In this study, we propose a novel imputation methodology for databases with non-response units by leveraging additional information from fully observed auxiliary variables. We assume that the variables included in the database are continuous and that the auxiliary variables, which are fully observed, help to improve the imputation capacity of the model. Within a fully Bayesian framework, our method utilizes a flexible mixture of multivariate normal distributions to jointly model the response and auxiliary variables. By employing the principles of Gaussian Cluster-Weighted modeling, we construct a predictive model to impute the missing values by leveraging information from the covariates. We present simulation studies and a real data illustration to demonstrate the imputation capacity of our method across various scenarios, comparing it to other methods in the literature

stat.ME↗

Imputation of Missing Data Using Linear Gaussian Cluster-Weighted Modeling

Missing data theory deals with the statistical methods in the occurrence of missing data. Missing data occurs when some values are not stored or observed for variables of interest. However, most of the statistical theory assumes that data is fully observed. An alternative to deal with incomplete databases is to fill in the spaces corresponding to the missing information based on some criteria, this technique is called imputation. We introduce a new imputation methodology for databases with univariate missing patterns based on additional information from fully-observed auxiliary variables. We assume that the non-observed variable is continuous, and that auxiliary variables assist to improve the imputation capacity of the model. In a fully Bayesian framework, our method uses a flexible mixture of multivariate normal distributions to model the response and the auxiliary variables jointly. Under this framework, we use the properties of Gaussian Cluster-Weighted modeling to construct a predictive model to impute the missing values using the information from the covariates. Simulations studies and a real data illustration are presented to show the method imputation capacity under a variety of scenarios and in comparison to other literature methods.

stat.ME↗

Alleviating Spatial Confounding in Spatial Frailty Models

Spatial confounding is how is called the confounding between fixed and spatial random effects. It has been widely studied and it gained attention in the past years in the spatial statistics literature, as it may generate unexpected results in modeling. The projection-based approach, also known as restricted models, appears as a good alternative to overcome the spatial confounding in generalized linear mixed models. However, when the support of fixed effects is different from the spatial effect one, this approach can no longer be applied directly. In this work, we introduce a method to alleviate the spatial confounding for the spatial frailty models family. This class of models can incorporate spatially structured effects and it is usual to observe more than one sample unit per area which means that the support of fixed and spatial effects differs. In this case, we introduce a two folded projection-based approach projecting the design matrix to the dimension of the space and then projecting the random effect to the orthogonal space of the new design matrix. To provide fast inference in our analysis we employ the integrated nested Laplace approximation methodology. The method is illustrated with an application with lung and bronchus cancer in California - US that confirms that the methodology efficiency.

stat.ME↗

Dynamic Time Scan Forecasting

The dynamic time scan forecasting method relies on the premise that the most important pattern in a time series precedes the forecasting window, i.e., the last observed values. Thus, a scan procedure is applied to identify similar patterns, or best matches, throughout the time series. As oppose to euclidean distance, or any distance function, a similarity function is dynamically estimated in order to match previous values to the last observed values. Goodness-of-fit statistics are used to find the best matches. Using the respective similarity functions, the observed values proceeding the best matches are used to create a forecasting pattern, as well as forecasting intervals. Remarkably, the proposed method outperformed statistical and machine learning approaches in a real case wind speed forecasting problem.

stat.AP↗

Inference on Dynamic Models for non-Gaussian Random Fields using INLA: A Homicide Rate Analysis of Brazilian Cities

Robust time series analysis is an important subject in statistical modeling. Models based on Gaussian distribution are sensitive to outliers, which may imply in a significant degradation in estimation performance as well as in prediction accuracy. State-space models, also referred as Dynamic Models, is a very useful way to describe the evolution of a time series variable through a structured latent evolution system. Integrated Nested Laplace Approximation (INLA) is a recent approach proposed to perform fast Bayesian inference in Latent Gaussian Models which naturally comprises Dynamic Models. We present how to perform fast and accurate non-Gaussian dynamic modeling with INLA and show how these models can provide a more robust time series analysis when compared with standard dynamic models based on Gaussian distributions. We formalize the framework used to fit complex non-Gaussian space-state models using the R package INLA and illustrate our approach in both a simulation study and on the brazilian homicide rate dataset.

stat.AP↗