SearcharxivSearch

arXiv subjects

Daniel Fryer

Publications and source records attributed to Daniel Fryer.

7 recordsLinked to original sources

Order selection with confidence for finite mixture models

The determination of the number of mixture components (the order) of a finite mixture model has been an enduring problem in statistical inference. We prove that the closed testing principle leads to a sequential testing procedure (STP) that allows for confidence statements to be made regarding the order of a finite mixture model. We construct finite sample tests, via data splitting and data swapping, for use in the STP, and we prove that such tests are consistent against fixed alternatives. Simulation studies and real data examples are used to demonstrate the performance of the finite sample tests-based STP, yielding practical recommendations of their use as confidence estimators in combination with point estimates such as the Akaike information or Bayesian information criteria. In addition, we demonstrate that a modification of the STP yields a method that consistently selects the order of a finite mixture model, in the asymptotic sense. Our STP is not only applicable for order selection of finite mixture models, but is also useful for making confidence statements regarding any sequence of nested models.

stat.ME

Shapley values for feature selection: The good, the bad, and the axioms

The Shapley value has become popular in the Explainable AI (XAI) literature, thanks, to a large extent, to a solid theoretical foundation, including four "favourable and fair" axioms for attribution in transferable utility games. The Shapley value is provably the only solution concept satisfying these axioms. In this paper, we introduce the Shapley value and draw attention to its recent uses as a feature selection tool. We call into question this use of the Shapley value, using simple, abstract "toy" counterexamples to illustrate that the axioms may work against the goals of feature selection. From this, we develop a number of insights that are then investigated in concrete simulation settings, with a variety of Shapley value formulations, including SHapley Additive exPlanations (SHAP) and Shapley Additive Global importancE (SAGE).

cs.LG

SARGDV: Efficient identification of groundwater-dependent vegetation using synthetic aperture radar

Groundwater depletion impacts the sustainability of numerous groundwater-dependent vegetation (GDV) globally, placing significant stress on their capacity to provide environmental and ecological support for flora, fauna, and anthropic benefits. Industries such as mining, agriculture, and plantations are heavily reliant on groundwater, the over-exploitation of which risks impacting groundwater regimes, quality, and accessibility for nearby GDVs. Cost effective methods of GDV identification will enable strategic protection of these critical ecological systems, through improved and sustainable groundwater management by communities and industry. Recent application of synthetic aperture radar (SAR) earth observation data in Australia has demonstrated the utility of radar for identifying terrestrial groundwater-dependent ecosystems at scale. We propose a robust classification method to advance identification of GDVs at scale using processed SAR data products adapted from a recent previous method. The method includes the development of SARGDV, a binary classification model, which uses the extreme gradient boosting (XGBoost) algorithm in conjunction with three data cubes composed of Sentinel-1 SAR interferometric wide images. The images were collected as a one-year time series over Mount Gambier, a region in South Australia, known to support GDVs. The SARGDV model demonstrated high performance for classifying GDVs with 77% precision, 76% true positive rate and 96% accuracy. This method may be used to support the protection of GDV communities globally by providing a long term, cost-effective solution to identify GDVs over variable regions and climates, via the use of freely available, high-resolution, globally available Sentinel-1 SAR data sets.

stat.AP

Shapley value confidence intervals for attributing variance explained

The coefficient of determination, the $R^2$, is often used to measure the variance explained by an affine combination of multiple explanatory covariates. An attribution of this explanatory contribution to each of the individual covariates is often sought in order to draw inference regarding the importance of each covariate with respect to the response phenomenon. A recent method for ascertaining such an attribution is via the game theoretic Shapley value decomposition of the coefficient of determination. Such a decomposition has the desirable efficiency, monotonicity, and equal treatment properties. Under a weak assumption that the joint distribution is pseudo-elliptical, we obtain the asymptotic normality of the Shapley values. We then utilize this result in order to construct confidence intervals and hypothesis tests regarding such quantities. Monte Carlo studies regarding our results are provided. We found that our asymptotic confidence intervals are computationally superior to competing bootstrap methods and are able to improve upon the performance of such intervals. In an expository application to Australian real estate price modelling, we employ Shapley value confidence intervals to identify significant differences between the explanatory contributions of covariates, between models, which otherwise share approximately the same $R^2$ value. These different models are based on real estate data from the same periods in 2019 and 2020, the latter covering the early stages of the arrival of the novel coronavirus, COVID-19.

stat.ME

Spherical data handling and analysis with R package rcosmo

The R package rcosmo was developed for handling and analysing Hierarchical Equal Area isoLatitude Pixelation(HEALPix) and Cosmic Microwave Background(CMB) radiation data. It has more than 100 functions. rcosmo was initially developed for CMB, but also can be used for other spherical data. This paper discusses transformations into rcosmo formats and handling of three types of non-CMB data: continuous geographic, point pattern and star-shaped. For each type of data we provide a brief description of the corresponding statistical model, data example and ready-to-use R code. Some statistical functionality of rcosmo is demonstrated for the example data converted into the HEALPix format. The paper can serve as the first practical guideline to transforming data into the HEALPix format and statistical analysis with rcosmo for geo-statisticians, GIS and R users and researches dealing with spherical data in non-HEALPix formats.

stat.AP

rcosmo: R Package for Analysis of Spherical, HEALPix and Cosmological Data

The analysis of spatial observations on a sphere is important in areas such as geosciences, physics and embryo research, just to name a few. The purpose of the package rcosmo is to conduct efficient information processing, visualisation, manipulation and spatial statistical analysis of Cosmic Microwave Background (CMB) radiation and other spherical data. The package was developed for spherical data stored in the Hierarchical Equal Area isoLatitude Pixelation (Healpix) representation. rcosmo has more than 100 different functions. Most of them initially were developed for CMB, but also can be used for other spherical data as rcosmo contains tools for transforming spherical data in cartesian and geographic coordinates into the HEALPix representation. We give a general description of the package and illustrate some important functionalities and benchmarks.

stat.CO