SearcharxivSearch

arXiv subjects

Nichole E. Carlson

Publications and source records attributed to Nichole E. Carlson.

2 recordsLinked to original sources

Adjusted Similarity Measures and a Violation of Expectations

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the property of 0 expectation under a null distribution and maximum value 1 under maximal similarity to aid in interpretation. Measures are frequently adjusted with respect to the permutation distribution for historic and analytic reasons. There is currently renewed interest in considering other null models more appropriate for context, such as clustering ensembles permitting a random number of identified clusters. The purpose of this work is two -- fold: (1) to generalize the study of the adjustment operator to general null models and to a more general procedure which includes statistical standardization as a special case and (2) to identify sufficient conditions for the adjustment operator to produce the intended properties, where sufficient conditions are related to whether and how observed data are incorporated into null distributions. We demonstrate how violations of the sufficient conditions may lead to substantial breakdown, such as by producing a non-positive measure under traditional adjustment rather than one with mean 0, or by producing a measure which is deterministically 0 under statistical standardization.

stat.ME

cpr: An R Package For Finding Parsimonious B-Spline Regression Models via Control Polygon Reduction and Control Net Reduction

The R package cpr provides tools for selection of parsimonious B-spline regression models via algorithms coined `control polygon reduction' (CPR) and `control net reduction' (CNR). B-Splines are commonly used in regression models to smooth data and approximate unknown functional forms. B-Splines are defined by a polynomial order and a knot sequence. Defining the knot sequence is non-trivial, but is critical with respect to the quality of the regression models. The focus of the CPR and CNR algorithms is to reduce a large knot sequence down to a parsimonious collection of elements while maintaining a high quality of fit. The algorithms are quick to implement and are flexible enough to support many types of data and regression approaches. The cpr package provides the end user collections of tools for the construction of B-spline basis matrices, construction of control polygons and control nets, and the use of diagnostics of the CPR and CNR algorithms.

stat.CO