SearcharxivSearch

arXiv subjects

Greg Kreider

Publications and source records attributed to Greg Kreider.

6 recordsLinked to original sources

Using Spacing to Detect Multi-Modality

The spacing of a unimodal variate resembles a `U' with a flat bottom and rapidly increasing values in the tails. When combined into multi-modal setups, the transition between variates into their tails forms local increases in the spacing. These peak features signal the data is multi-modal while the flats locate the modes. We develop tests of the features, including parametric models based on an assumed null univariate distribution, the number of runs or longest run in the signed difference of the interval spacing, and feature reconstruction by means of permutations of the runs or bootstrap samples from a pool based on the difference of the signal. We also try combining existing changepoint detectors to look where the behavior of the spacing changes. This paper describes the processing of the spacing and models and tests in full, gives examples of their use, and summarizes the performance in general.

stat.ME

Evaluating Spacing Tests of Multi-Modality

The spacing of data contains information about the underlying modality. Sections of consistent, similar spacing correspond to modes while it increases between them. We have defined parametric, runs-based non-parametric, and data driven tests to evaluate these features --- flats and peaks --- and determine the presence and location of multiple modes. This report will check the tests in two situations. By varying a bi-modal setup we can control the changes in spacing and determine the resolution and sensitivity of the tests. By applying them to the large number of test cases that exist in the literature from other modality studies we check the stability and consistency of the results. We will also evaluate the accuracy of the null-distribution models of the features. The results show that the spacing does reflect the data's modality, that the tests do screen marginal cases, and where the analysis begins to break down.

stat.ME

Runs and Bootstrap Tests For Signal Feature Significance

Runs tests have long been used as a non-parametric check if data contains a non-random signal. We derive a recursive expression for the distribution of the longest run using Markov chain theory. Next we develop a permutation test on the runs comprising a feature to get the probability of its height. This leads finally to a bootstrap test on the height using the raw, continuous data. Such a test can evaluate not only the large heights of peaks but also the small heights of flats. We can apply these tests to features in the spacing of data to detect and locate multi-modality.

stat.ME

Modality Analysis via Spacing with the Dimodal Software Libraries

Spacing, the difference between consecutive order statistics, has two features that reflect the modality of the data. Consistent, stable values occur around modes while local increases mark the transitions between them. These features not only signal multi-modality, they also locate modes and anti-modes. Dimodal is an R package for detecting and evaluating these situations. It includes parametric feature models and bootstrap tests for spacing smoothed by low-pass filtering, non-parametric runs and permutation tests for the interval spacing, and a fusion of changepoints in the raw spacing. We introduce the analysis, describe the package, its implementation and performance, and apply it to identifying Kirkwood gaps in the asteroid belt. We also present ports of the software, with DimodalCPy a command-line program written in C with a Python interface.

stat.ME

A Spacing Estimator

The distribution of the spacing, or the difference between consecutive order statistics, is known only for uniform and exponential random variates. We add here logistic and Gumbel variates, and present an estimator for distributions with a known inverse cumulative density function. We show the estimator is accurate to the limit of numerical simulations for points near the middle of the order statistics, but degrades by up to 20% in the tails.

stat.ME

Interval Spacing

We define interval spacing as the difference in the order statistics of data over a gap of some width. We derive its density, expected value, and variance for uniform, exponential, and logistic variates. We show that interval spacing is equivalent to running a rectangular low-pass filter over the spacing, which simplifies the expressions for the expected values and introduces correlations between overlapping intervals.

stat.ME