SearcharxivSearch

arXiv subjects

Charles D. Coleman

Publications and source records attributed to Charles D. Coleman.

6 recordsLinked to original sources

A New Way to Look at Regional Survey Data: Differences in Vacancy Rates and Persons per Household by County, 2000-2005

Regional survey estimates and their significance levels are simultaneously displayed in maps that show all 3,141 U.S. counties and equivalents. An analyst can focus his attention on significant differences (or those with a different, low-valued uncertainty measure) for all but the very smallest counties. Differences between Census 2000 and the 2005 American Community Survey values are shown.

stat.ME

The Asymptotic Equivalence of Level-Based and Share-Based Loss Functions

Level-based and share-based loss functions are asymptotically equivalent if, in the limit, their averages converge almost surely to a constant ratio. These loss functions take a target value and its realization as arguments and are often used to measure accuracy. The equivalence is proved for a large class of loss functions, the weighted exponentiated functions, when the weights are decomposable as a particular product form. An upshot is that when losses are averaged for a large number of units, differences in ratios and, hence, ranks, are negligible, when the average (or summed) difference between the target values and their realizations is around zero. This implies the almost sure asymptotic convergence of numerical and distributive accuracy when using these loss functions.

math.ST

Loss Functions for Detecting Outliers in Panel Data

The detection of outliers is of critical importance in the assurance of data quality. Outliers may exist in observed data or in data derived from these observed data, such as estimates and forecasts. An outlier may indicate a problem with its data generation process or may simply be a true, but unusual, statement about the world. Without making any distributional assumptions, we proposes the use of loss functions to detect these outliers in panel data. Part I covers nonnegative data. We axiomatically derive an unsigned loss function. We then develop a signed loss function ito account for positive and negative outliers separately. In the case of nominal time we obtain an exact parametrization of the loss function. A time-invariant loss function permits the comparison of data at multiple times on the same basis. We provide several examples, including an example in which the outliers are classified by another variable. Part II covers data of mixed sign. Similar to Part I, we axiomatically develop unsigned and signed loss functions. We search for optimal values of the loss function parameter using graphs.

stat.ME

Loss Functions for Measuring the Accuracy of Nonnegative Cross-Sectional Predictions

Measuring the accuracy of cross-sectional predictions is a subjective problem. Generally, this problem is avoided. In contrast, this paper confronts subjectivity up front by eliciting an impartial decision-maker's preferences. These preferences are embedded into an axiomatically-derived loss function, one of the simplest version of which is described. The parameters of the loss function can be estimated by linear regression. Specification tests for this function are described. This framework is extended to weighted averages of estimates to find the optimal weightings. A special case occurs when the predictions represent resource allocations: the apportionment literature is used to construct the Webster-Saint Lagüe Rule, a particular parametrization of the loss function. These loss functions are compared to those existing in the literature. Finally, a family of bias measures are created using signed versions of these loss functions.

stat.ME

Total Loss Functions for Measuring the Accuracy of Nonnegative Cross-Sectional Predictions

The total loss function associated with a set of cross-sectional predictions, that is, estimates or forecasts, summarizes the set's overall accuracy. Its arguments are the individual cross-sectional units' loss functions. Under general assumptions, including impartiality, about the forms of the individual loss functions, and the specific assumptions that the total loss function is anonymous and monotonic, only the additive, multiplicative and L-type (with restrictions) total loss functions are found to be admissible. The first two total loss functions correspond to different interpretations of economic utility. An isomorphism exists between these two total loss functions. Thus, the additive total loss function can always be used. This isomorphism can also be used to explore the properties of various combinations of total and individual loss functions. Moreover, the additive loss function obeys the von Neumann-Morgenstern expected utility axioms.

stat.ME

The Importance of Variable Importance

Variable importance is defined as a measure of each regressor's contribution to model fit. Using R^2 as the fit criterion in linear models leads to the Shapley value (LMG) and proportionate value (PMVD) as variable importance measures. Similar measures are defined for ensemble models, using random forests as the example. The properties of the LMG and PMVD are compared. Variable importance is proposed to assess regressors' practical effects or "oomph." The uses of variable importance in modelling, interventions and causal analysis are discussed.

stat.ME