SearcharxivSearch

arXiv subjects

Ronaldo Dias

Publications and source records attributed to Ronaldo Dias.

16 recordsLinked to original sources

Bayesian Hierarchical Modeling for Predicting Spatially Correlated Curves in Irregular Domains: A Case Study on PM10 Pollution

This study presents a Bayesian hierarchical model for analyzing spatially correlated functional data and handling irregularly spaced observations. The model uses Bernstein polynomial (BP) bases combined with autoregressive random effects, allowing for nuanced modeling of spatial correlations between sites and dependencies of observations within curves. Moreover, the proposed procedure introduces a distinct structure for the random effect component compared to previous works. Simulation studies conducted under various challenging scenarios verify the model's robustness, demonstrating its capacity to accurately recover spatially dependent curves and predict observations at unmonitored locations. The model's performance is further supported by its application to real-world data, specifically PM$_{10}$ particulate matter measurements from a monitoring network in Mexico City. This application is of practical importance, as particles can penetrate the respiratory system and aggravate various health conditions. The model effectively predicts concentrations at unmonitored sites, with uncertainty estimates that reflect spatial variability across the domain. This new methodology provides a flexible framework for the FDA in spatial contexts and addresses challenges in analyzing irregular domains with potential applications in environmental monitoring.

stat.ME

Variational Inference for Bayesian Bridge Regression

We study the implementation of Automatic Differentiation Variational inference (ADVI) for Bayesian inference on regression models with bridge penalization. The bridge approach uses $\ell_α$ norm, with $α\in (0, +\infty)$ to define a penalization on large values of the regression coefficients, which includes the Lasso ($α= 1$) and ridge $(α= 2)$ penalizations as special cases. Full Bayesian inference seamlessly provides joint uncertainty estimates for all model parameters. Although MCMC aproaches are available for bridge regression, it can be slow for large dataset, specially in high dimensions. The ADVI implementation allows the use of small batches of data at each iteration (due to stochastic gradient based algorithms), therefore speeding up computational time in comparison with MCMC. We illustrate the approach on non-parametric regression models with B-splines, although the method works seamlessly for other choices of basis functions. A simulation study shows the main properties of the proposed method.

stat.ML

Wavelet estimation of nonstationary spatial covariance function

This work proposes a new procedure for estimating the non-stationary spatial covariance function for Spatial-Temporal Deformation. The proposed procedure is based on a monotonic function approach. The deformation functions are expanded as a linear combination of the wavelet basis. The estimate of the deformation guarantees an injective transformation. Such that two distinct locations in the geographic plane are not mapped into the same point in the deformation plane. Simulation studies have shown the effectiveness of this procedure. An application to historical daily maximum temperature records exemplifies the flexibility of the proposed methodology when dealing with real datasets.

stat.ME

Bayesian Variable Selection for Function-on-Scalar Regression Models: a comparative analysis

In this work, we developed a new Bayesian method for variable selection in function-on-scalar regression (FOSR). Our method uses a hierarchical Bayesian structure and latent variables to enable an adaptive covariate selection process for FOSR. Extensive simulation studies show the proposed method's main properties, such as its accuracy in estimating the coefficients and high capacity to select variables correctly. Furthermore, we conducted a substantial comparative analysis with the main competing methods, the BGLSS (Bayesian Group Lasso with Spike and Slab prior) method, the group LASSO (Least Absolute Shrinkage and Selection Operator), the group MCP (Minimax Concave Penalty), and the group SCAD (Smoothly Clipped Absolute Deviation). Our results demonstrate that the proposed methodology is superior in correctly selecting covariates compared with the existing competing methods while maintaining a satisfactory level of goodness of fit. In contrast, the competing methods could not balance selection accuracy with goodness of fit. We also considered a COVID-19 dataset and some socioeconomic data from Brazil as an application and obtained satisfactory results. In short, the proposed Bayesian variable selection model is highly competitive, showing significant predictive and selective quality.

stat.ME

Clustering Functional Data via Variational Inference

Functional data analysis deals with data recorded densely over time (or any other continuum) with one or more observed curves per subject. Conceptually, functional data are continuously defined, but in practice, they are usually observed at discrete points. Among different kinds of functional data analyses, clustering analysis aims to determine underlying groups of curves in the dataset when there is no information on the group membership of each individual curve. In this work, we propose a new model-based approach for clustering and smoothing functional data simultaneously via variational inference. We derive coordinate ascent mean-field variational Bayes algorithms to approximate the posterior distribution of our model parameters by finding the variational distribution with the smallest Kullback-Leibler divergence to the posterior. The performance of our proposed method is evaluated using simulated data and publicly available datasets.

stat.ME

Bayesian Adaptive Selection of Basis Functions for Functional Data Representation

Considering the context of functional data analysis, we developed and applied a new Bayesian approach via Gibbs sampler to select basis functions for a finite representation of functional data. The proposed methodology uses Bernoulli latent variables to assign zero to some of the basis function coefficients with a positive probability. This procedure allows for an adaptive basis selection since it can determine the number of bases and which should be selected to represent functional data. Moreover, the proposed procedure measures the uncertainty of the selection process and can be applied to multiple curves simultaneously. The methodology developed can deal with observed curves that may differ due to experimental error and random individual differences between subjects, which one can observe in a real dataset application involving daily numbers of COVID-19 cases in Brazil. Simulation studies show the main properties of the proposed method, such as its accuracy in estimating the coefficients and the strength of the procedure to find the true set of basis functions. Despite having been developed in the context of functional data analysis, we also compared the proposed model via simulation with the well-established LASSO and Bayesian LASSO, which are methods developed for non-functional data.

stat.ME

Modeling the Evolution of Infectious Diseases with Functional Data Models: The Case of COVID-19 in Brazil

In this paper, we apply statistical methods for functional data to explain the heterogeneity in the evolution of number of deaths of Covid-19 over different regions. We treat the cumulative daily number of deaths in a specific region as a curve (functional data) such that the data comprise of a set of curves over a cross-section of locations. We start by using clustering methods for functional data to identify potential heterogeneity in the curves and their functional derivatives. This first stage is an unconditional descriptive analysis, as we do not use any covariate to estimate the clusters. The estimated clusters are analyzed as "levels of alert" to identify cities in a possible critical situation. In the second and final stage, we propose a functional quantile regression model of the death curves on a number of scalar socioeconomic and demographic indicators in order to investigate their functional effects at different levels of the cumulative number of deaths over time. The proposed model showed a superior predictive capacity by providing better curve fit at different levels of the cumulative number of deaths compared to the functional regression model based on ordinary least squares.

stat.AP

Variational Full Bayes Lasso: Knots Selection in Regression Splines

We develop a fully automatic Bayesian Lasso via variational inference. This is a scalable procedure for approximating the posterior distribution. Special attention is driven to the knot selection in regression spline. In order to carry through our proposal, a full automatic variational Bayesian Lasso, a Jefferey's prior is proposed for the hyperparameters and a decision theoretical approach is introduced to decide if a knot is selected or not. Extensive simulation studies were developed to ensure the effectiveness of the proposed algorithms. The performance of the algorithms were also tested in some real data sets, including data from the world pandemic Covid-19. Again, the algorithms showed a very good performance in capturing the data structure.

stat.ME

A Basis Approach to Surface Clustering

This paper presents a novel method for clustering surfaces. The proposal involves first using basis functions in a tensor product to smooth the data and thus reduce the dimension to a finite number of coefficients, and then using these estimated coefficients to cluster the surfaces via the k-means algorithm. An extension of the algorithm to clustering tensors is also discussed. We show that the proposed algorithm exhibits the property of strong consistency, with or without measurement errors, in correctly clustering the data as the sample size increases. Simulation studies suggest that the proposed method outperforms the benchmark k-means algorithm which uses the original vectorized data. In addition, an EGG real data example is considered to illustrate the practical application of the proposal.

stat.ME

Scalable modeling of nonstationary covariance functions with non-folding B-spline deformation

We propose a method for nonstationary covariance function modeling, based on the spatial deformation method of Sampson and Guttorp [1992], but using a low-rank, scalable deformation function written as a linear combination of the tensor product of B-spline basis. This approach addresses two important weaknesses in current computational aspects. First, it allows one to constrain estimated 2D deformations to be non-folding (bijective) in 2D. This requirement of the model has, up to now,been addressed only by arbitrary levels of spatial smoothing. Second, basis functions with compact support enable the application to large datasets of spatial monitoring sites of environmental data. An application to rainfall data in southeastern Brazil illustrates the method

stat.ME

Selection of the Number of Clusters in Functional Data Analysis

Identifying the number $K$ of clusters in a dataset is one of the most difficult problems in clustering analysis. A choice of $K$ that correctly characterizes the features of the data is essential for building meaningful clusters. In this paper we tackle the problem of estimating the number of clusters in functional data analysis by introducing a new measure that can be used with different procedures in selecting the optimal $K$. The main idea is to use a combination of two test statistics, which measure the lack of parallelism and the mean distance between curves, to compute criteria such as the within and between cluster sum of squares. Simulations in challenging scenarios suggest that procedures using this measure can detect the correct number of clusters more frequently than existing methods in the literature. The application of the proposed method is illustrated on several real datasets.

stat.ME

Analysis of Aggregated Functional Data from Mixed Populations with Application to Energy Consumption

Understanding the energy consumption patterns of different types of consumers is essential in any planning of energy distribution. However, obtaining consumption information for single individuals is often either not possible or too expensive. Therefore, we consider data from aggregations of energy use, that is, from sums of individuals' energy use, where each individual falls into one of C consumer classes. Unfortunately, the exact number of individuals of each class may be unknown: consumers do not always report the appropriate class, due to various factors including differential energy rates for different consumer classes. We develop a methodology to estimate the expected energy use of each class as a function of time and the true number of consumers in each class. We also provide some measure of uncertainty of the resulting estimates. To accomplish this, we assume that the expected consumption is a function of time that can be well approximated by a linear combination of B-splines. Individual consumer perturbations from this baseline are modeled as B-splines with random coefficients. We treat the reported numbers of consumers in each category as random variables with distribution depending on the true number of consumers in each class and on the probabilities of a consumer in one class reporting as another class. We obtain maximum likelihood estimates of all parameters via a maximization algorithm. We introduce a special numerical trick for calculating the maximum likelihood estimates of the true number of consumers in each class. We apply our method to a data set and study our method via simulation.

stat.AP

Aggregated functional data model for Near-Infrared Spectroscopy calibration and prediction

Calibration and prediction for NIR spectroscopy data are performed based on a functional interpretation of the Beer-Lambert formula. Considering that, for each chemical sample, the resulting spectrum is a continuous curve obtained as the summation of overlapped absorption spectra from each analyte plus a Gaussian error, we assume that each individual spectrum can be expanded as a linear combination of B-splines basis. Calibration is then performed using two procedures for estimating the individual analytes curves: basis smoothing and smoothing splines. Prediction is done by minimizing the square error of prediction. To assess the variance of the predicted values, we use a leave-one-out jackknife technique. Departures from the standard error models are discussed through a simulation study, in particular, how correlated errors impact on the calibration step and consequently on the analytes' concentration prediction. Finally, the performance of our methodology is demonstrated through the analysis of two publicly available datasets.

stat.ME

Genetic Algorithm for Constrained Optimization with Stochastic Feasibility Region with Application to Vehicle Path Planning

In real-time trajectory planning for unmanned vehicles, on-board sensors, radars and other instruments are used to collect information on possible obstacles to be avoided and pathways to be followed. Since, in practice, observations of the sensors have measurement errors, the stochasticity of the data has to be incorporated into the models. In this paper, we consider using a genetic algorithm for the constrained optimization problem of finding the trajectory with minimum length between two locations, avoiding the obstacles on the way. To incorporate the variability of the sensor readings, we propose a more general framework, where the feasible regions of the genetic algorithm are stochastic. In this way, the probability that a possible solution of the search space, say x, is feasible can be derived from the random observations of obstacles and pathways, creating a real-time data learning algorithm. By building a confidence region from the observed data such that its border intersects with the solution point x, the level of the confidence region defines the probability that x is feasible. We propose using a smooth penalty function based on the Gaussian distribution, facilitating the borders of the feasible regions to be reached by the algorithm.

stat.ME

A Review of Kernel Density Estimation with Applications to Econometrics

Nonparametric density estimation is of great importance when econometricians want to model the probabilistic or stochastic structure of a data set. This comprehensive review summarizes the most important theoretical aspects of kernel density estimation and provides an extensive description of classical and modern data analytic methods to compute the smoothing parameter. Throughout the text, several references can be found to the most up-to-date and cut point research approaches in this area, while econometric data sets are analyzed as examples. Lastly, we present SIZer, a new approach introduced by Chaudhuri and Marron (2000), whose objective is to analyze the visible features representing important underlying structures for different bandwidths.

stat.ME

A Hierarchical Model for Aggregated Functional Data

In many areas of science one aims to estimate latent sub-population mean curves based only on observations of aggregated population curves. By aggregated curves we mean linear combination of functional data that cannot be observed individually. We assume that several aggregated curves with linear independent coefficients are available. More specifically, we assume each aggregated curve is an independent partial realization of a Gaussian process with mean modeled through a weighted linear combination of the disaggregated curves. We model the mean of the Gaussian processes as a smooth function approximated by a function belonging to a finite dimensional space ${\cal H}_K$ which is spanned by $K$ B-splines basis functions. We explore two different specifications of the covariance function of the Gaussian process: one that assumes a constant variance across the domain of the process, and a more general variance structure which is itself modelled as a smooth function, providing a nonstationary covariance function. Inference procedure is performed following the Bayesian paradigm allowing experts' opinion to be considered when estimating the disaggregated curves. Moreover, it naturally provides the uncertainty associated with the parameters estimates and fitted values. Our model is suitable for a wide range of applications. We concentrate on two different real examples: calibration problem for NIR spectroscopy data and an analysis of distribution of energy among different type of consumers.

stat.ME