SearcharxivSearch

arXiv subjects

Ronny Vallejos

Publications and source records attributed to Ronny Vallejos.

11 recordsLinked to original sources

Welcome to the Statverse: A Metaverse for Data Science

This paper introduces the Statverse, a Metaverse framework designed to revolutionize statistical education in the digital age. Our key goal is to report our progress and encourage others to integrate similar strategies into their programs. The proposed framework seamlessly integrates the physical and digital realms to provide an immersive environment for the nuanced representation of complex statistical concepts. Finally, we discuss the potential impact of Statverse on advancing Statistical Education, offering a transformative approach to teaching and learning in the digital age. Statverse is the outcome of an academic partnership between Universidad Técnica Federico Santa María (UTFSM) and the University of Edinburgh (UoE).

stat.OT

Agreement coefficients for continuous variables: A review

Agreement coefficients provide a fundamental framework for quantifying the concordance between two or more measurement methods applied to the same continuous variable. Unlike correlation, which measures the strength of a linear relationship, agreement focuses on assessing whether measurements are numerically similar, capturing both precision and accuracy. This review provides a comprehensive overview of the primary statistical approaches for assessing agreement between continuous variables. Such a synthesis is timely, as it has been 15-20 years since the last major review in the field. Beginning with the seminal contributions of Bland and Altman (1986) and Lin (1989), the paper discusses extensions of their methods to robust, multivariate, and repeated-measures settings, as well as recent developments like the probability of agreement and measures based on alternative distance functions measures. Special attention is given to probabilistic and spatial generalizations, including frameworks designed for geostatistical and areal data, which have become increasingly relevant in modern applications such as image analysis and environmental statistics. Through illustrative examples and comparative discussions, this review highlights the evolution, connections, and limitations of existing agreement measures, identifying open challenges and directions for future research.

stat.ME

A Kolmogorov-Arnold Neural Model for Cascading Extremes

This paper addresses the growing concern of cascading extreme events, such as an extreme earthquake followed by a tsunami, by presenting a novel method for risk assessment focused on these domino effects. The proposed approach develops an extreme value theory framework within a Kolmogorov-Arnold network (KAN) to estimate the probability of one extreme event triggering another, conditionally on a feature vector. An extra layer is added to the KAN architecture to ensure that the parameter of interest lies within the unit interval, and we refer to the resulting neural model as KANE (KAN with Natural Enforcement). The proposed method is backed by exhaustive numerical studies and further illustrated with real-world applications to seismology and climatology.

stat.ME

Effective Sample Size for Functional Spatial Data

The effective sample size quantifies the amount of independent information contained in a dataset, accounting for redundancy due to correlation between observations. While widely used in geostatistics for scalar data, its extension to functional spatial data has remained largely unexplored. In this work, we introduce a novel definition of the effective sample size for functional geostatistical data, employing the trace-covariogram as a measure of correlation, and show that it retains the intuitive properties of the classical scalar ESS. We illustrate the behavior of this measure using a functional autoregressive process, demonstrating how serial dependence and the allocation of variability across eigen-directions influence the resulting functional ESS. Finally, the approach is applied to a real meteorological dataset of geometric vertical velocities over a portion of the Earth, showing how the method can quantify redundancy and determine the effective number of independent curves in functional spatial datasets.

stat.ME

Optimized imaging prefiltering for enhanced image segmentation

The Box-Cox transformation, introduced in 1964, is a widely used statistical tool for stabilizing variance and improving normality in data analysis. Its application in image processing, particularly for image enhancement, has gained increasing attention in recent years. This paper investigates the use of the Box-Cox transformation as a preprocessing step for image segmentation, with a focus on the estimation of the transformation parameter. We evaluate the effectiveness of the transformation by comparing various segmentation methods, highlighting its advantages for traditional machine learning techniques-especially in situations where no training data is available. The results demonstrate that the transformation enhances feature separability and computational efficiency, making it particularly beneficial for models like discriminant analysis. In contrast, deep learning models did not show consistent improvements, underscoring the differing impacts of the transformation across model types and image characteristics.

stat.AP

A new coefficient to measure agreement between continuous variables

Assessing agreement between two instruments is crucial in clinical studies to evaluate the similarity between two methods measuring the same subjects. This paper introduces a novel coefficient, termed rho1, to measure agreement between continuous variables, focusing on scenarios where two instruments measure experimental units in a study. Unlike existing coefficients, rho1 is based on L1 distances, making it robust to outliers and not relying on nuisance parameters. The coefficient is derived for bivariate normal and elliptically contoured distributions, showcasing its versatility. In the case of normal distributions, rho1 is linked to Lin's coefficient, providing a useful alternative. The paper includes theoretical properties, an inference framework, and numerical experiments to validate the performance of rho1. This novel coefficient presents a valuable tool for researchers assessing agreement between continuous variables in various fields, including clinical studies and spatial analysis.

stat.ME

A concordance coefficient for lattice data: An application to poverty indices in Chile

This paper introduces a novel coefficient for measuring agreement between two lattice sequences observed in the same areal units, motivated by the analysis of different methodologies for measuring poverty rates in Chile. Building on the multivariate concordance coefficient framework, our approach accounts for dependencies in the multivariate lattice process using a non-negative definite matrix of weights, assuming a Multivariate Conditionally Autoregressive (GMCAR) process. We adopt a Bayesian perspective for inference, using summaries from Bayesian estimates. The methodology is illustrated through an analysis of poverty rates in the Metropolitan and Valparaíso regions of Chile, with High Posterior Density (HPD) intervals provided for the poverty rates. This work addresses a methodological gap in the understanding of agreement coefficients and enhances the usability of these measures in the context of social variables typically assessed in areal units.

stat.ME

Comparing two spatial variables with the probability of agreement

Computing the agreement between two continuous sequences is of great interest in statistics when comparing two instruments or one instrument with a gold standard. The probability of agreement (PA) quantifies the similarity between two variables of interest, and it is useful for accounting what constitutes a practically important difference. In this article we introduce a generalization of the PA for the treatment of spatial variables. Our proposal makes the PA dependent on the spatial lag. As a consequence, for isotropic stationary and nonstationary spatial processes, the conditions for which the PA decays as a function of the distance lag are established. Estimation is addressed through a first-order approximation that guarantees the asymptotic normality of the sample version of the PA. The sensitivity of the PA is studied for finite sample size, with respect to the covariance parameters. The new method is described and illustrated with real data involving autumnal changes in the green chromatic coordinate (Gcc), an index of "greenness" that captures the phenological stage of tree leaves, is associated with carbon flux from ecosystems, and is estimated from repeated images of forest canopies.

stat.ME

A Spatial Concordance Correlation Coefficient with an Application to Image Analysis

In this work we define a spatial concordance coefficient for second-order stationary processes. This problem has been widely addressed in a non-spatial context, but here we consider a coefficient that for a fixed spatial lag allows one to compare two spatial sequences along a 45-degree line. The proposed coefficient was explored for the bivariate Matérn and Wendland covariance functions. The asymptotic normality of a sample version of the spatial concordance coefficient for an increasing domain sampling framework was established for the Wendland covariance function. To work with large digital images, we developed a local approach for estimating the concordance that uses local spatial models on non-overlapping windows. Monte Carlo simulations were used to gain additional insights into the asymptotic properties for finite sample sizes. As an illustrative example, we applied this methodology to two similar images of a deciduous forest canopy. The images were recorded with different cameras but similar fields-of-view and within minutes of each other. Our analysis showed that the local approach helped to explain a percentage of the non-spatial concordance and to provided additional information about its decay as a function of the spatial lag.

stat.ME

Sensitivity of codispersion to noise and error in ecological and environmental data

Codispersion analysis is a new statistical method developed to assess spatial covariation between two spatial processes that may not be isotropic or stationary. Its application to anisotropic ecological datasets have provided new insights into mechanisms underlying observed patterns of species distributions and the relationship between individual species and underlying environmental gradients. However, the performance of the codispersion coefficient when there is noise or measurement error ("contamination") in the data has been addressed only theoretically. Here, we use Monte Carlo simulations and real datasets to investigate the sensitivity of codispersion to four types of contamination commonly seen in many real-world environmental and ecological studies. Three of these involved examining codispersion of a spatial dataset with a contaminated version of itself. The fourth examined differences in codisperson between plants and soil conditions, where the estimates of soil characteristics were based on complete or thinned datasets. In all cases, we found that estimates of codispersion were robust when contamination, such as data thinning, was relatively low (<15\%), but were sensitive to larger percentages of contamination. We also present a useful method for imputing missing spatial data and discuss several aspects of the codispersion coefficient when applied to noisy data to gain more insight about the performance of codispersion in practice.

stat.ME

SpatialPack: Computing the Association Between Two Spatial Processes

An R package SpatialPack that implements routines to compute point estimators and perform hypothesis testing of the spatial association between two stochastic sequences is introduced. These methods address the spatial association between two processes that have been observed over the same spatial locations. We briefly review the methodologies for which the routines are developed. The core routines have been implemented in C and linked to R to ensure a reasonable computational speed. Three examples are presented to illustrate the use of the package with both simulated and real data. The particular case of computing the association between two time series is also considered. Besides elementary plots and outputs we also provide a plot to visualize the spatial correlation in all directions using a new graphical tool called codispersion map. The potential extensions of SpatialPack are also discussed.

stat.AP