SearcharxivSearch

arXiv subjects

Karel Hron

Publications and source records attributed to Karel Hron.

14 recordsLinked to original sources

Dimension reduction of multivariate densities in Bayes spaces

The Bayes space provides a Hilbert space structure for analysing probability density functions (PDFs), equipping them with a geometry that reflects their relative and constrained nature. A key tool in this framework is the centred logratio (clr) transformation, which establishes an isometric isomorphism between the Bayes space and (a subspace of) the classical $L^2$ space. This makes it possible to apply functional data analysis (FDA) techniques, particularly functional principal component analysis (FPCA), to both univariate and multivariate density data in the context of dimension reduction. For multivariate PDFs, embedding them in the Bayes space enables an orthogonal decomposition into independent and interactive components. Furthermore, the independent part can be decomposed into mutually orthogonal geometric marginals. This structure provides more profound insights into the sources of variation in multivariate densities. We show that this decomposition of the total variance is optimal in a PCA sense, impacting the interpretation of the eigenfunctions and scores resulting from FPCA. We demonstrate that applying FPCA directly to multivariate densities is equivalent in a certain sense to applying multivariate FPCA to their decomposed form, with the resulting eigenfunctions and scores decomposing accordingly. The unique decomposition based on these theoretical results is applied to housing and geological empirical data respectively, demonstrating the interpretability and practical value of this approach.

stat.ME

Compositional Periodic Spline Approximation for Circular Density Data in Bayes Spaces

This paper proposes a novel framework for the approximation and analysis of circular density data using compositional periodic splines within Bayes spaces with the Hilbert space structure. By applying the centered log-ratio transformation, densities are represented in a subspace of the standard $L^2$ space of real-valued functions, which enables the use of functional data analysis tools while preserving the relative nature of distributions and their periodic structure. A coefficient-based construction of periodic splines with a zero-integral constraint is developed, together with matrix formulations for both smoothing splines and penalized splines, allowing efficient estimation and implementation. The methodology is applied to long-term wind direction data, where it provides smooth and interpretable density estimates and supports further statistical analysis, including functional regression. The results demonstrate the practical relevance of the proposed approach and its potential for extensions to more complex density-valued data.

stat.ME

Robust functional PCA for relative data

This paper introduces a robust approach to functional principal component analysis (FPCA) for relative data, particularly density functions. While recent papers have studied density data within the Bayes space framework, there has been limited focus on developing robust methods to effectively handle anomalous observations and large noise. To address this, we extend the Mahalanobis distance concept to Bayes spaces, proposing its regularized version that accounts for the constraints inherent in density data. Based on this extension, we introduce a new method, robust density principal component analysis (RDPCA), for more accurate estimation of functional principal components in the presence of outliers. The method's performance is validated through simulations and real-world applications, showing its ability to improve covariance estimation and principal component analysis compared to traditional methods.

stat.ME

Approximation of bivariate densities with compositional splines

Reliable estimation and approximation of probability density functions is fundamental for their further processing. However, their specific properties, i.e. scale invariance and relative scale, prevent the use of standard methods of spline approximation and have to be considered when building a suitable spline basis. Bayes Hilbert space methodology allows to account for these properties of densities and enables their conversion to a standard Lebesgue space of square integrable functions using the centered log-ratio transformation. As the transformed densities fulfill a zero integral constraint, the constraint should likewise be respected by any spline basis used. Bayes Hilbert space methodology also allows to decompose bivariate densities into their interactive and independent parts with univariate marginals. As this yields a useful framework for studying the dependence structure between random variables, a spline basis ideally should admit a corresponding decomposition. This paper proposes a new spline basis for (transformed) bivariate densities respecting the desired zero integral property. We show that there is a one-to-one correspondence of this basis to a corresponding basis in the Bayes Hilbert space of bivariate densities using tools of this methodology. Furthermore, the spline representation and the resulting decomposition into interactive and independent parts are derived. Finally, this novel spline representation is evaluated in a simulation study and applied to empirical geochemical data.

stat.ME

Efficient spline orthogonal basis for representation of density functions

Probability density functions form a specific class of functional data objects with intrinsic properties of scale invariance and relative scale characterized by the unit integral constraint. The Bayes spaces methodology respects their specific nature, and the centred log-ratio transformation enables processing such functional data in the standard Lebesgue space of square-integrable functions. As the data representing densities are frequently observed in their discrete form, the focus has been on their spline representation. Therefore, the crucial step in the approximation is to construct a proper spline basis reflecting their specific properties. Since the centred log-ratio transformation forms a subspace of functions with a zero integral constraint, the standard $B$-spline basis is no longer suitable. Recently, a new spline basis incorporating this zero integral property, called $Z\!B$-splines, was developed. However, this basis does not possess the orthogonal property which is beneficial from computational and application point of view. As a result of this paper, we describe an efficient method for constructing an orthogonal $Z\!B$-splines basis, called $Z\!B$-splinets. The advantages of the $Z\!B$-splinet approach are foremost a computational efficiency and locality of basis supports that is desirable for data interpretability, e.g. in the context of functional principal component analysis. The proposed approach is demonstrated on an empirical demographic dataset.

stat.ME

Identifying Important Pairwise Logratios in Compositional Data with Sparse Principal Component Analysis

Compositional data are characterized by the fact that their elemental information is contained in simple pairwise logratios of the parts that constitute the composition. While pairwise logratios are typically easy to interpret, the number of possible pairs to consider quickly becomes (too) large even for medium-sized compositions, which might hinder interpretability in further multivariate analyses. Sparse methods can therefore be useful to identify few, important pairwise logratios (respectively parts contained in them) from the total candidate set. To this end, we propose a procedure based on the construction of all possible pairwise logratios and employ sparse principal component analysis to identify important pairwise logratios. The performance of the procedure is demonstrated both with simulated and real-world data. In our empirical analyses, we propose three visual tools showing (i) the balance between sparsity and explained variability, (ii) stability of the pairwise logratios, and (iii) importance of the original compositional parts to aid practitioners with their model interpretation.

stat.ME

Exploratory functional data analysis of multivariate densities for the identification of agricultural soil contamination by risk elements

Geochemical mapping of risk element concentrations in soils is performed in countries around the world. It results in large datasets of high analytical quality, which can be used to identify soils that violate individual legislative limits for safe food production. However, there is a lack of advanced data mining tools that would be suitable for sensitive exploratory data analysis of big data while respecting the natural variability of soil composition. To distinguish anthropogenic contamination from natural variation, the analysis of the entire data distributions for smaller sub-areas is key. In this article, we propose a new data mining method for geochemical mapping data based on functional data analysis of probability density functions in the framework of Bayes spaces after post-stratification of a big dataset to smaller districts. Proposed tools allow us to analyse the entire distribution, going beyond a superficial detection of extreme concentration anomalies. We illustrate the proposed methodology on a dataset gathered according to the Czech national legislation (1990--2009). Taking into account specific properties of probability density functions and recent results for orthogonal decomposition of multivariate densities enabled us to reveal real contamination patterns that were so far only suspected in Czech agricultural soils. We process the above Czech soil composition dataset by first compartmentalising it into spatial units, in particular the districts, and by subsequently clustering these districts according to diagnostic features of their uni- and multivariate distributions at high concentration ends. Comparison between compartments is key to the reliable distinction of diffuse contamination. In this work, we used soil contamination by Cu-bearing pesticides as an example for empirical testing of the proposed data mining approach.

stat.AP

Orthogonal decomposition of multivariate densities in Bayes spaces and its connection with copulas

Bayes spaces were initially designed to provide a geometric framework for the modeling and analysis of distributional data. It has recently come to light that this methodology can be exploited to provide an orthogonal decomposition of bivariate probability distributions into an independent and an interaction part. In this paper, new insights into these results are provided by reformulating them using Hilbert space theory and a multivariate extension is developed using a distributional analog of the Hoeffding-Sobol identity. A connection between the resulting decomposition of a multivariate density and its copula-based representation is also provided.

math.ST

Compositional Cubes: A New Concept for Multi-factorial Compositions

Compositional data are commonly known as multivariate observations carrying relative information. Even though the case of vector or even two-factorial compositional data (compositional tables) is already well described in the literature, there is still a need for a comprehensive approach to the analysis of multi-factorial relative-valued data. Therefore, this contribution builds around the current knowledge about compositional data a general theory of work with k-factorial compositional data. As a main finding it turns out that similar to the case of compositional tables also the multi-factorial structures can be orthogonally decomposed into an independent and several interactive parts and, moreover, a coordinate representation allowing for their separate analysis by standard analytical methods can be constructed. For the sake of simplicity, these features are explained in detail for the case of three-factorial compositions (compositional cubes), followed by an outline covering the general case. The three-dimensional structure is analysed in depth in two practical examples, dealing with systems of spatial and time dependent compositional cubes. The methodology is implemented in the R package robCompositions.

stat.ME

Bivariate Densities in Bayes Spaces: Orthogonal Decomposition and Spline Representation

A new orthogonal decomposition for bivariate probability densities embedded in Bayes Hilbert spaces is derived. It allows one to represent a density into independent and interactive parts, the former being built as the product of revised definitions of marginal densities and the latter capturing the dependence between the two random variables being studied. The developed framework opens new perspectives for dependence modelling (which is commonly performed through copulas), and allows for the analysis of dataset of bivariate densities, in a Functional Data Analysis perspective. A spline representation for bivariate densities is also proposed, providing a computational cornerstone for the developed theory.

math.ST

Compositional splines for representation of density functions

In the context of functional data analysis, probability density functions as non-negative functions are characterized by specific properties of scale invariance and relative scale which enable to represent them with the unit integral constraint without loss of information. On the other hand, all these properties are a challenge when the densities need to be approximated with spline functions, including construction of the respective spline basis. The Bayes space methodology of density functions enables to express them as real functions in the standard $L^2$ space using the centered log-ratio transformation. The resulting functions satisfy the zero integral constraint. This is a key to propose a new spline basis, holding the same property, and consequently to build a new class of spline functions, called compositional splines, which can approximate probability density functions in a consistent way. The paper provides also construction of smoothing compositional splines and possible orthonormalization of the spline basis which might be useful in some applications. Finally, statistical processing of densities using the new approximation tool is demonstrated in case of simplicial functional principal component analysis with anthropometric data.

math.NA

Robust Principal Component Analysis for Compositional Tables

A data table which is arranged according to two factors can often be considered as a compositional table. An example is the number of unemployed people, split according to gender and age classes. Analyzed as compositions, the relevant information would consist of ratios between different cells of such a table. This is particularly useful when analyzing several compositional tables jointly, where the absolute numbers are in very different ranges, e.g. if unemployment data are considered from different countries. Within the framework of the logratio methodology, compositional tables can be decomposed into independent and interactive parts, and orthonormal coordinates can be assigned to these parts. However, these coordinates usually require some prior knowledge about the data, and they are not easy to handle for exploring the relationships between the given factors. Here we propose a special choice of coordinates with a direct relation to centered logratio (clr) coefficients, which are particularly useful for an interpretation in terms of the original cells of the tables. With these coordinates, robust principal component analysis (PCA) is performed for dimension reduction, allowing to investigate the relationships between the factors. The link between orthonormal coordinates and clr coefficients enables to apply robust PCA, which would otherwise suffer from the singularity of clr coefficients.

stat.ME

Interpretation of Compositional Regression with Application to Time Budget Analysis

Regression with compositional response or covariates, or even regression between parts of a composition, is frequently employed in social sciences. Among other possible applications, it may help to reveal interesting features in time allocation analysis. As individual activities represent relative contributions to the total amount of time, statistical processing of raw data (frequently represented directly as proportions or percentages) using standard methods may lead to biased results. Specific geometrical features of time budget variables are captured by the logratio methodology of compositional data, whose aim is to build (preferably orthonormal) coordinates to be applied with popular statistical methods. The aim of this paper is to present recent tools of regression analysis within the logratio methodology and apply them to reveal potential relationships among psychometric indicators in a real-world data set. In particular, orthogonal logratio coordinates have been introduced to enhance the interpretability of coefficients in regression models.

math.ST

Preprocessing of centred logratio transformed density functions using smoothing splines

With large-scale database systems, statistical analysis of data, formed by probability distributions, become an important task in explorative data analysis. Nevertheless, due to specific properties of density functions, their proper statistical treatment still represents a challenging task in functional data analysis. Namely, the usual L2 metric does not fully accounts for the relative character of information, carried by density functions; instead, their geometrical features are followed by Bayes spaces of measures. The easiest possibility of expressing density functions in L2 space is to use centred logratio transformation, nevertheless, it results in functional data with a constant integral constraint that needs to be taken into account for further analysis. While theoretical background for reasonable analysis of density functions is already provided comprehensively by Bayes spaces themselves, preprocessing issues still need to be developed. The aim of this paper is to introduce optimal smoothing splines for centred logratio transformed density functions that take all their specific features into account and provide a concise methodology for reasonable preprocessing of raw (discretized) distributional observations. Theoretical developments are illustrated with a real-world data set from official statistics.

math.NA