SearcharxivSearch

arXiv subjects

Robert L. Obenchain

Publications and source records attributed to Robert L. Obenchain.

10 recordsLinked to original sources

The Efficient Shrinkage Path: Maximum Likelihood of Minimum MSE Risk

A new generalized ridge regression shrinkage path is proposed that is as short as possible under the restriction that it must pass through the vector of regression coefficient estimators that make the overall Optimal Variance-Bias Trade-Off under Normal distribution-theory. Five distinct types of ridge TRACE displays plus other graphics for this efficient path are motivated and illustrated here. These visualizations provide invaluable data-analytic insights and improved self-confidence to researchers and data scientists fitting linear models to ill-conditioned (confounded) data.

stat.ME

Incremental Cost-Effectiveness Statistical Inference: Calculations and Communications

We illustrate use of nonparametric statistical methods to compare alternative treatments for a particular disease or condition on both their relative effectiveness and their relative cost. These Incremental Cost Effectiveness (ICE) methods are based upon Bootstrapping, i.e. Resampling with Replacement from observational or clinical-trial data on individual patients. We first show how a reasonable numerical value for the "Shadow Price of Health" can be chosen using functions within the ICEinfer R-package when effectiveness is not measured in "QALY"s. We also argue that simple histograms are ideal for communicating key findings to regulators, while our more detailed graphics may well be more informative and compelling for other health-care stakeholders.

stat.ME

Nonlinear Generalized Ridge Regression

A Two-Stage approach is described that literally "straighten outs" any potentially nonlinear relationship between a y-outcome variable and each of p = 2 or more potential x-predictor variables. The y-outcome is then predicted from all p of these "linearized" spline-predictors using the form of Generalized Ridge Regression that is most likely to yield minimal MSE risk under Normal distribution-theory. These estimates are then compared and contrasted with those from the Generalized Additive Model that uses the same x-variables.

stat.ME

Nonparametric Generalized Ridge Regression

A Two-Stage approach enables researchers to make optimal non-linear predictions via Generalized Ridge Regression using models that contain two or more x-predictor variables and make only realistic minimal assumptions. The optimal regression coefficient estimates that result are either unbiased or most likely to have mininal MSE risk under Normal distribution theory. All necessary calculations and graphical displays are generated using current versions of CRAN R-packages. A numerical example using the "corrected" USArrests data.frame introduces and illustrates this new robust statistical methodology. While applying this strategy to regression models with several hundred observations is straight-forward, the computations required in such cases can be extensive.

stat.ME

Mortality Rates of US Counties: Are they Reliable and Predictable?

We examine US County-level observational data on Lung Cancer mortality rates in 2012 and overall Circulatory Respiratory mortality rates in 2016 as well as their "Top Ten" potential causes from Federal or State sources. We find that these two mortality rates for 2,812 US Counties have remarkably little in common. Thus, for predictive modeling, we use a single "compromise" measure of mortality that has several advantages. The vast majority of our new findings have simple implications that we illustrate graphically.

stat.AP

EPA Particulate Matter Data -- Analyses using Local Control Strategy

Statistical Learning methodology for analysis of large collections of cross-sectional observational data can be most effective when the approach used is both Nonparametric and Unsupervised. We illustrate use of our NU Learning approach on 2016 US environmental epidemiology data that we have made freely available. We encourage other researchers to download these data, apply whatever methodology they wish, and contribute to development of a broad-based ``consensus view'' of potential effects of Secondary Organic Aerosols (volatile organic compounds of predominantly biogenic or anthropogenic origin) within PM2.5 particulate matter on circulatory and/or respiratory mortality. Our analyses here focus on the question: ``Are regions with relatively high air-borne biogenic particulate matter also expected to have relatively high circulatory and/or respiratory mortality?''

cs.CY

Maximum Likelihood Ridge Regression

My first paper exclusively about ridge regression was published in Technometrics and chosen for invited presentation at the 1975 Joint Statistical Meetings in Atlanta. Unfortunately, that paper contained a wide range of assorted details and results. Luckily, Gary McDonald's published discussion of that paper focused primarily on my use of Maximum Likelihood estimation under normal distribution-theory. In this review of some results from all four of my ridge publications between 1975 and 2022, I highlight the Maximum Likelihood findings that appear to be most important in practical application of shrinkage in regression.

stat.ME

Ridge TRACE Diagnostics

We describe a new p-parameter generalized ridge-regression shrinkage-pattern recently implemented in the RXshrink CRAN R-package. The five distinct types of ridge TRACE displays discussed and illustrated here provide invaluable data-analytic insights and improved self-confidence to researchers and data scientists fitting linear models to ill-conditioned datasets.

stat.ME

Affine Reduction of Dimensionality: An Origin-Centric Perspective

We consider statistical methods for reduction of multivariate dimensionality that have invariance and/or commutativity properties under the affine group of transformations (origin translations plus linear combinations of coordinates along initial axes). The methods discussed here differ from traditional principal component and coordinate approaches in that they are origin-centric. Because all Cartesian coordinates of the origin are zero, it is the unique fixed point for subsequent linear transformations of point scatters. Whenever visualizations allow shifting between and/or combining of Cartesian and polar coordinate representations, as in Biplots, the location of this origin is critical. Specifically, origin-centric visualizations enhance the psychology of graphical perception by yielding scatters that can be interpreted as Dyson swarms. The key factor is typically the analyst's choice of origin via an initial "centering" translation; this choice determines whether the recovered scatter will have either no points depicted as being near the origin or else one (or more) points exactly coincident with this origin.

stat.ME

Bias and response heterogeneity in an air quality data set

It is well-known that claims coming from observational studies often fail to replicate when rigorously re-tested. The technical problems include multiple testing, multiple modeling and bias. Any or all of these problems can give rise to claims that will fail to replicate. There is a need for statistical methods that are easily applied, are easy to understand, and are likely to give reliable results. In particular, simple ways for reducing the influence of bias are essential. In this paper, the Local Control method developed by Robert Obenchain is explicated using a small air quality/longevity data set first analyzed in the New England Journal of Medicine. The benefits of our paper are twofold. First, we describe a reliable strategy for analysis of observational data. Second and importantly, the global claim that longevity increases with improvements in air quality made in the NEJM paper needs to be modified. There is subgroup heterogeneity in the effect of air quality on longevity (one size does not fit all), and this heterogeneity is largely explained by factors other than air quality.

stat.AP