SearcharxivSearch

arXiv subjects

Michael Law

Publications and source records attributed to Michael Law.

10 recordsLinked to original sources

Distributional Robustness and Transfer Learning Through Empirical Bayes

We consider the problem of statistical inference on parameters of a target population when auxiliary observations are available from related populations. We propose a flexible empirical Bayes approach that can be applied on top of any asymptotically linear estimator to incorporate information from related populations when constructing confidence regions. The proposed methodology is valid regardless of whether there are direct observations on the population of interest. We demonstrate the performance of the empirical Bayes confidence regions on synthetic data as well as on the Trends in International Mathematics and Sciences Study when using the debiased Lasso as the basic algorithm in high-dimensional regression.

math.ST

Invariant Probabilistic Prediction

In recent years, there has been a growing interest in statistical methods that exhibit robust performance under distribution changes between training and test data. While most of the related research focuses on point predictions with the squared error loss, this article turns the focus towards probabilistic predictions, which aim to comprehensively quantify the uncertainty of an outcome variable given covariates. Within a causality-inspired framework, we investigate the invariance and robustness of probabilistic predictions with respect to proper scoring rules. We show that arbitrary distribution shifts do not, in general, admit invariant and robust probabilistic predictions, in contrast to the setting of point prediction. We illustrate how to choose evaluation metrics and restrict the class of distribution shifts to allow for identifiability and invariance in the prototypical Gaussian heteroscedastic linear model. Motivated by these findings, we propose a method to yield invariant probabilistic predictions, called IPP, and study the consistency of the underlying parameters. Finally, we demonstrate the empirical performance of our proposed procedure on simulated as well as on single-cell data.

stat.ME

Longitudinal Position and Cancer Risk in the United States Revisited

Background: The debate over daylight saving time has surged, with interests in the effects of sunlight exposure on health. \commentnj{Prior studies simulated daylight saving time and standard time conditions by analyzing different locations within time zones and neighboring areas across time zone borders. Methods: We analyzed cancer incidence rates from various longitudinal positions within time zones and at time zone borders in the contiguous United States. Using data from State Cancer Profiles (2016-2020), we analyzed total cancer of 19 types and specific rates for eight cancers, adjusted for age and includes all demographics. Log-linear regression is used to replicate a previous study, and spatial regression models are employed to explore discontinuities at borders. Results: Cancer rate differences lack statistical significance within time zones and near borders for total cancer and most individual cancers. Exceptions included breast, prostate, and liver \& bile duct cancers, which exhibited significant relationships with relative position at the 95\% significance level. Breast and liver and bile duct cancers saw decreases, while prostate cancer incidence increased from west to east within time zones. Conclusions: Relative position does not have a significant impact on cancer incidence, hence cancer development in general. Isolated exceptions may warrant further investigation as more data becomes available. Impact: Our findings challenge prior research, revealing numerous inconsistencies. These disparities urge a reconsideration of the potential disparities in human health associated with daylight saving time and standard time. They offer insights contribute to the ongoing discussion surrounding the retention or abandonment of DST.

stat.AP

A Rank-Based Sequential Test of Independence

We consider the problem of independence testing for two univariate random variables in a sequential setting. By leveraging recent developments on safe, anytime-valid inference, we propose a test with time-uniform type I error control and derive explicit bounds on the finite sample performance of the test. We demonstrate the empirical performance of the procedure in comparison to existing sequential and non-sequential independence tests. Furthermore, since the proposed test is distribution free under the null hypothesis, we empirically simulate the gap due to Ville's inequality, the supermartingale analogue of Markov's inequality, that is commonly applied to control type I error in anytime-valid inference, and apply this to construct a truncated sequential test.

stat.ME

Large-scale detector testing for the GAPS Si(Li) Tracker

Lithium-drifted silicon [Si(Li)] has been used for decades as an ionizing radiation detector in nuclear, particle, and astrophysical experiments, though such detectors have frequently been limited to small sizes (few cm$^2$) and cryogenic operating temperatures. The 10-cm-diameter Si(Li) detectors developed for the General Antiparticle Spectrometer (GAPS) balloon-borne dark matter experiment are novel particularly for their requirements of low cost, large sensitive area (~10 m$^2$ for the full 1440-detector array), high temperatures (near -40$\,^\circ$C), and energy resolution below 4 keV FWHM for 20--100-keV x-rays. Previous works have discussed the manufacturing, passivation, and small-scale testing of prototype GAPS Si(Li) detectors. Here we show for the first time the results from detailed characterization of over 1100 flight detectors, illustrating the consistent intrinsic low-noise performance of a large sample of GAPS detectors. This work demonstrates the feasibility of large-area and low-cost Si(Li) detector arrays for next-generation astrophysics and nuclear physics applications.

physics.ins-det

Rank-Constrained Least-Squares: Prediction and Inference

In this work, we focus on the high-dimensional trace regression model with a low-rank coefficient matrix. We establish a nearly optimal in-sample prediction risk bound for the rank-constrained least-squares estimator under no assumptions on the design matrix. Lying at the heart of the proof is a covering number bound for the family of projection operators corresponding to the subspaces spanned by the design. By leveraging this complexity result, we perform a power analysis for a permutation test on the existence of a low-rank signal under the high-dimensional trace regression model. We show that the permutation test based on the rank-constrained least-squares estimator achieves non-trivial power with no assumptions on the minimum (restricted) eigenvalue of the covariance matrix of the design. Finally, we use alternating minimization to approximately solve the rank-constrained least-squares problem to evaluate its empirical in-sample prediction risk and power of the resulting permutation test in our numerical study.

math.ST

High-Dimensional Varying Coefficient Models with Functional Random Effects

We consider a sparse high-dimensional varying coefficients model with random effects, a flexible linear model allowing covariates and coefficients to have a functional dependence with time. For each individual, we observe discretely sampled responses and covariates as a function of time as well as time invariant covariates. Under sampling times that are either fixed and common or random and independent amongst individuals, we propose a projection procedure for the empirical estimation of all varying coefficients. We extend this estimator to construct confidence bands for a fixed number of varying coefficients.

math.ST

The Hot Hand and Its Effect on the NBA

This paper aims to revisit and expand upon previous work on the "hot hand" phenomenon in basketball, specifically in the NBA. Using larger, modern data sets, we test streakiness of shooting patterns and the presence of hot hand behavior in free throw shooting, while going further by examining league-wide hot hand trends and the changes in individual player behavior. Additionally, we perform simulations in order to assess their power. While we find no evidence of the hot hand in game-play and only weak evidence in free throw trials, we find that some NBA players exhibit behavioral changes based on the outcome of their previous shot.

stat.AP

Inference Without Compatibility

We consider hypotheses testing problems for three parameters in high-dimensional linear models with minimal sparsity assumptions of their type but without any compatibility conditions. Under this framework, we construct the first $\sqrt{n}$-consistent estimators for low-dimensional coefficients, the signal strength, and the noise level. We support our results using numerical simulations and provide comparisons with other estimators.

math.ST

Estimating the Random Effect in Big Data Mixed Models

We consider three problems in high-dimensional Gaussian linear mixed models. Without any assumptions on the design for the fixed effects, we construct an asymptotic $F$-statistic for testing whether a collection of random effects is zero, derive an asymptotic confidence interval for a single random effect at the parametric rate $\sqrt{n}$, and propose an empirical Bayes estimator for a part of the mean vector in ANOVA type models that performs asymptotically as well as the oracle Bayes estimator. We support our results with numerical simulations and provide comparisons with oracle estimators. The procedures developed are applied to the Trends in International Mathematics and Sciences Study (TIMSS) data.

math.ST