SearcharxivSearch

arXiv subjects

Lawrence Brown

Publications and source records attributed to Lawrence Brown.

9 recordsLinked to original sources

On the Energy Dependence of Galactic Cosmic Ray Anisotropies in the Very Local Interstellar Medium

We report on the energy dependence of galactic cosmic rays (GCRs) in the very local interstellar medium (VLISM) as measured by the Low Energy Charged Particle (LECP) instrument on the Voyager 1 (V1) spacecraft. The LECP instrument includes a dual-ended solid state detector particle telescope mechanically scanning through 360 deg across eight equally-spaced angular sectors. As reported previously, LECP measurements showed a dramatic increase in GCR intensities for all sectors of the >=211 MeV count rate (CH31) at the V1 heliopause (HP) crossing in 2012, however, since then the count rate data have demonstrated systematic episodes of intensity decrease for particles around 90{\deg} pitch angle. To shed light on the energy dependence of these GCR anisotropies over a wide range of energies, we use V1 LECP count rate and pulse height analyzer (PHA) data from >=211 MeV channel together with lower energy LECP channels. Our analysis shows that while GCR anisotropies are present over a wide range of energies, there is a decreasing trend in the amplitude of second-order anisotropy with increasing energy during anisotropy episodes. A stronger pitch-angle scattering at the higher velocities is argued as a potential cause for this energy dependence. A possible cause for this velocity dependence arising from weak rigidity dependence of the scattering mean free path and resulting velocity-dominated scattering rate is discussed. This interpretation is consistent with a recently reported lack of corresponding GCR electron anisotropies.

astro-ph.HE

Models as Approximations I: Consequences Illustrated with Linear Regression

In the early 1980s Halbert White inaugurated a "model-robust'' form of statistical inference based on the "sandwich estimator'' of standard error. This estimator is known to be "heteroskedasticity-consistent", but it is less well-known to be "nonlinearity-consistent'' as well. Nonlinearity, however, raises fundamental issues because in its presence regressors are not ancillary, hence can't be treated as fixed. The consequences are deep: (1)~population slopes need to be re-interpreted as statistical functionals obtained from OLS fits to largely arbitrary joint $\xy$~distributions; (2)~the meaning of slope parameters needs to be rethought; (3)~the regressor distribution affects the slope parameters; (4)~randomness of the regressors becomes a source of sampling variability in slope estimates; (5)~inference needs to be based on model-robust standard errors, including sandwich estimators or the $\xy$~bootstrap. In theory, model-robust and model-trusting standard errors can deviate by arbitrary magnitudes either way. In practice, significant deviations between them can be detected with a diagnostic test.

stat.ME

Models as Approximations II: A Model-Free Theory of Parametric Regression

We develop a model-free theory of general types of parametric regression for iid observations. The theory replaces the parameters of parametric models with statistical functionals, to be called "regression functionals'', defined on large non-parametric classes of joint $\xy$ distributions, without assuming a correct model. Parametric models are reduced to heuristics to suggest plausible objective functions. An example of a regression functional is the vector of slopes of linear equations fitted by OLS to largely arbitrary $\xy$ distributions, without assuming a linear model (see Part~I). More generally, regression functionals can be defined by minimizing objective functions or solving estimating equations at joint $\xy$ distributions. In this framework it is possible to achieve the following: (1)~define a notion of well-specification for regression functionals that replaces the notion of correct specification of models, (2)~propose a well-specification diagnostic for regression functionals based on reweighting distributions and data, (3)~decompose sampling variability of regression functionals into two sources, one due to the conditional response distribution and another due to the regressor distribution interacting with misspecification, both of order $N^{-1/2}$, (4)~exhibit plug-in/sandwich estimators of standard error as limit cases of $\xy$ bootstrap estimators, and (5)~provide theoretical heuristics to indicate that $\xy$ bootstrap standard errors may generally be more stable than sandwich estimators.

math.ST

Assumption Lean Regression

It is well known that models used in conventional regression analysis are commonly misspecified. A standard response is little more than a shrug. Data analysts invoke Box's maxim that all models are wrong and then proceed as if the results are useful nevertheless. In this paper, we provide an alternative. Regression models are treated explicitly as approximations of a true response surface that can have a number of desirable statistical properties, including estimates that are asymptotically unbiased. Valid statistical inference follows. We generalize the formulation to include regression functionals, which broadens substantially the range of potential applications. An empirical application is provided to illustrate the paper's key concepts.

stat.ME

Calibrated Percentile Double Bootstrap For Robust Linear Regression Inference

We consider inference for the parameters of a linear model when the covariates are random and the relationship between response and covariates is possibly non-linear. Conventional inference methods such as z-intervals perform poorly in these cases. We propose a double bootstrap-based calibrated percentile method, perc-cal, as a general-purpose CI method which performs very well relative to alternative methods in challenging situations such as these. The superior performance of perc-cal is demonstrated by a thorough, full-factorial design synthetic data study as well as a real data example involving the length of criminal sentences. We also provide theoretical justification for the perc-cal method under mild conditions. The method is implemented in the R package `perccal', available through CRAN and coded primarily in C++, to make it easier for practitioners to use.

stat.ME

Optimal shrinkage estimation of mean parameters in family of distributions with quadratic variance

This paper discusses the simultaneous inference of mean parameters in a family of distributions with quadratic variance function. We first introduce a class of semiparametric/parametric shrinkage estimators and establish their asymptotic optimality properties. Two specific cases, the location-scale family and the natural exponential family with quadratic variance function, are then studied in detail. We conduct a comprehensive simulation study to compare the performance of the proposed methods with existing shrinkage estimators. We also apply the method to real data and obtain encouraging results.

math.ST

Improved Precision in Estimating Average Treatment Effects

The Average Treatment Effect (ATE) is a global measure of the effectiveness of an experimental treatment intervention. Classical methods of its estimation either ignore relevant covariates or do not fully exploit them. Moreover, past work has considered covariates as fixed. We present a method for improving the precision of the ATE estimate: the treatment and control responses are estimated via a regression, and information is pooled between the groups to produce an asymptotically unbiased estimate; we subsequently justify the random X paradigm underlying the result. Standard errors are derived, and the estimator's performance is compared to the traditional estimator. Conditions under which the regression-based estimator is preferable are detailed, and a demonstration on real data is presented.

stat.ME

Valid post-selection inference

It is common practice in statistical data analysis to perform data-driven variable selection and derive statistical inference from the resulting model. Such inference enjoys none of the guarantees that classical statistical theory provides for tests and confidence intervals when the model has been chosen a priori. We propose to produce valid ``post-selection inference'' by reducing the problem to one of simultaneous inference and hence suitably widening conventional confidence and retention intervals. Simultaneity is required for all linear functions that arise as coefficient estimates in all submodels. By purchasing ``simultaneity insurance'' for all possible submodels, the resulting post-selection inference is rendered universally valid under all possible model selection procedures. This inference is therefore generally conservative for particular selection procedures, but it is always less conservative than full Scheffe protection. Importantly it does not depend on the truth of the selected submodel, and hence it produces valid inference even in wrong models. We describe the structure of the simultaneous inference problem and give some asymptotic results.

math.ST

Alternative formulas for synthetic dual system estimation in the 2000 census

The U.S. Census Bureau provides an estimate of the true population as a supplement to the basic census numbers. This estimate is constructed from data in a post-censal survey. The overall procedure is referred to as dual system estimation. Dual system estimation is designed to produce revised estimates at all levels of geography, via a synthetic estimation procedure. We design three alternative formulas for dual system estimation and investigate the differences in area estimates produced as a result of using those formulas. The primary target of this exercise is to better understand the nature of the homogeneity assumptions involved in dual system estimation and their consequences when used for the enumeration data that occurs in an actual large scale application like the Census. (Assumptions of this nature are sometimes collectively referred to as the ``synthetic assumption'' for dual system estimation.) The specific focus of our study is the treatment of the category of census counts referred to as imputations in dual system estimation. Our results show the degree to which varying treatment of these imputation counts can result in differences in population estimates for local areas such as states or counties.

stat.AP