Searcharxiv⌕ Search

arXiv subjects

Ya'acov Ritov

Publications and source records attributed to Ya'acov Ritov.

At least 19 recordsLinked to original sources

Empirical Bayes Estimation of the Mean of a Function of the Latent Variable with Applications to the Treatment of Nonresponse

We consider the estimation of linear functionals of the mixing distribution in a nonparametric empirical Bayes framework. Our main interest is in situations in which the mixing distribution is only partially identifiable, as may arise in complex sampling situations with nonresponse. We argue that estimating the functional by applying it to the semiparametric maximum likelihood estimator of the mixing distribution is an efficient tool, even when the maximum likelihood estimator is not unique.

math.ST↗

The Root Finding Problem Revisited: Beyond the Robbins-Monro procedure

We introduce Sequential Probability Ratio Bisection (SPRB), a novel stochastic approximation algorithm that adapts to the local behavior of the (regression) function of interest around its root. We establish theoretical guarantees for SPRB's asymptotic performance, showing that it achieves the optimal convergence rate and minimal asymptotic variance even when the target function's derivative at the root is small (at most half the step size), a regime where the classical Robbins-Monro procedure typically suffers reduced convergence rates. Further, we show that if the regression function is discontinuous at the root, Robbins-Monro converges at a rate of $1/n$ whilst SPRB attains exponential convergence. If the regression function has vanishing first-order derivative, SPRB attains a faster rate of convergence compared to stochastic approximation. As part of our analysis, we derive a nonasymptotic bound on the expected sample size and establish a generalized Central Limit Theorem under random stopping times. Remarkably, SPRB automatically provides nonasymptotic time-uniform confidence sequences that do not explicitly require knowledge of the convergence rate. We demonstrate the practical effectiveness of SPRB through simulation results.

math.ST↗

Limitations of refinement methods for weak to strong generalization

Standard techniques for aligning large language models (LLMs) utilize human-produced data, which could limit the capability of any aligned LLM to human level. Label refinement and weak training have emerged as promising strategies to address this superalignment problem. In this work, we adopt probabilistic assumptions commonly used to study label refinement and analyze whether refinement can be outperformed by alternative approaches, including computationally intractable oracle methods. We show that both weak training and label refinement suffer from irreducible error, leaving a performance gap between label refinement and the oracle. These results motivate future research into developing alternative methods for weak to strong generalization that synthesize the practicality of label refinement or weak training and the optimality of the oracle procedure.

stat.ML↗

From Thomas Bayes to Big Data: On the feasibility of being a subjective Bayesian

We argue that the Bayesian paradigm, of a prior which represents the beliefs of the statistician before observing the data, is not feasible in ultra-high-dimensional models. We claim that natural priors that represent the a priori beliefs fail in unpredictable ways under values of the parameters that cannot be honestly ignored. We do not claim that the frequentist estimators we present cannot be mimicked by Bayesian procedures, but that these Bayesian procedures do not represent beliefs. They were created with the frequentist analysis in mind, and in most cases, they cannot represent a consistent set of beliefs about the parameters (for example, since they depend on the loss function, the particular functional of interest, and not only on the a priori knowledge, different priors should be used for different analyses of the same data set). In a way, these are frequentist procedures using a Bayesian technique. The paper presents different examples where the subjective point of view fails. It is argued that the arguments based on Wald's and Savage's seminal works are not relevant to the validity of the subjective Bayesian paradigm. The discussion tries to deal with the fundamentals, but the argument is based on a firm mathematical proofs.

math.ST↗

Estimation and Inference for the Average Treatment Effect in a Score-Explained Heterogeneous Treatment Effect Model

In many practical situations, randomly assigning treatments to subjects is uncommon due to feasibility constraints. For example, economic aid programs and merit-based scholarships are often restricted to those meeting specific income or exam score thresholds. In these scenarios, traditional approaches to estimating treatment effects typically focus solely on observations near the cutoff point, thereby excluding a significant portion of the sample and potentially leading to information loss. Moreover, these methods generally achieve a non-parametric convergence rate. While some approaches, e.g., Mukherjee et al. (2021), attempt to tackle these issues, they commonly assume that treatment effects are constant across individuals, an assumption that is often unrealistic in practice. In this study, we propose a differencing and matching-based estimator of the average treatment effect on the treated (ATT) in the presence of heterogeneous treatment effects, utilizing all available observations. We establish the asymptotic normality of our estimator and illustrate its effectiveness through various synthetic and real data analyses. Additionally, we demonstrate that our method yields non-parametric estimates of the conditional average treatment effect (CATE) and individual treatment effect (ITE) as a byproduct.

stat.ME↗

A transfer learning framework for weak-to-strong generalization

Modern large language model (LLM) alignment techniques rely on human feedback, but it is unclear whether these techniques fundamentally limit the capabilities of aligned LLMs. In particular, it is unknown if it is possible to align (stronger) LLMs with superhuman capabilities with (weaker) human feedback without degrading their capabilities. This is an instance of the weak-to-strong generalization problem: using feedback from a weaker (less capable) model to train a stronger (more capable) model. We prove that weak-to-strong generalization is possible by eliciting latent knowledge from pre-trained LLMs. In particular, we cast the weak-to-strong generalization problem as a transfer learning problem in which we wish to transfer a latent concept prior from a weak model to a strong pre-trained model. We prove that a naive fine-tuning approach suffers from fundamental limitations, but an alternative refinement-based approach suggested by the problem structure provably overcomes the limitations of fine-tuning. Finally, we demonstrate the practical applicability of the refinement approach in multiple LLM alignment tasks.

stat.ML↗

A mixture of a normal distribution with random mean and variance -- Examples of inconsistency of maximum likelihood estimates

We consider the estimation of the mixing distribution of a normal distribution where both the shift and scale are unobserved random variables. We argue that in general, the model is not identifiable. We give an elegant non-constructive proof that the model is identifiable if the shift parameter is bounded by a known value. However, we argue that the generalized maximum likelihood estimator is inconsistent even if the shift parameter is bounded and the shift and scale parameters are independent. The mixing distribution, however, is identifiable if we have more than one observations per any realization of the latent shift and scale.

math.ST↗

No need for an oracle: the nonparametric maximum likelihood decision in the compound decision problem is minimax

We discuss the asymptotics of the nonparametric maximum likelihood estimator (NPMLE) in the normal mixture model. We then prove the convergence rate of the NPMLE decision in the empirical Bayes problem with normal observations. We point to (and heavily use) the connection between the NPMLE decision and Stein unbiased risk estimator (\sure). Next, we prove that the same solution is optimal in the compound decision problem where the unobserved parameters are not assumed to be random. Similar results are usually claimed using an oracle-based argument. However, we contend that the standard oracle argument is not valid. It was only partially proved that it can be fixed, and the existing proofs of these partial results are tedious. Our approach, on the other hand, is straightforward and short.

math.ST↗

Algorithmic Fairness in Performative Policy Learning: Escaping the Impossibility of Group Fairness

In many prediction problems, the predictive model affects the distribution of the prediction target. This phenomenon is known as performativity and is often caused by the behavior of individuals with vested interests in the outcome of the predictive model. Although performativity is generally problematic because it manifests as distribution shifts, we develop algorithmic fairness practices that leverage performativity to achieve stronger group fairness guarantees in social classification problems (compared to what is achievable in non-performative settings). In particular, we leverage the policymaker's ability to steer the population to remedy inequities in the long term. A crucial benefit of this approach is that it is possible to resolve the incompatibilities between conflicting group fairness definitions.

stat.ML↗

Learning In Reverse Causal Strategic Environments With Ramifications on Two Sided Markets

Motivated by equilibrium models of labor markets, we develop a formulation of causal strategic classification in which strategic agents can directly manipulate their outcomes. As an application, we compare employers that anticipate the strategic response of a labor force with employers that do not. We show through a combination of theory and experiment that employers with performatively optimal hiring policies improve employer reward, labor force skill level, and in some cases labor force equity. On the other hand, we demonstrate that performative employers harm labor force utility and fail to prevent discrimination in other cases.

stat.ML↗

Distributional Robustness and Transfer Learning Through Empirical Bayes

We consider the problem of statistical inference on parameters of a target population when auxiliary observations are available from related populations. We propose a flexible empirical Bayes approach that can be applied on top of any asymptotically linear estimator to incorporate information from related populations when constructing confidence regions. The proposed methodology is valid regardless of whether there are direct observations on the population of interest. We demonstrate the performance of the empirical Bayes confidence regions on synthetic data as well as on the Trends in International Mathematics and Sciences Study when using the debiased Lasso as the basic algorithm in high-dimensional regression.

math.ST↗

Longitudinal Position and Cancer Risk in the United States Revisited

Background: The debate over daylight saving time has surged, with interests in the effects of sunlight exposure on health. \commentnj{Prior studies simulated daylight saving time and standard time conditions by analyzing different locations within time zones and neighboring areas across time zone borders. Methods: We analyzed cancer incidence rates from various longitudinal positions within time zones and at time zone borders in the contiguous United States. Using data from State Cancer Profiles (2016-2020), we analyzed total cancer of 19 types and specific rates for eight cancers, adjusted for age and includes all demographics. Log-linear regression is used to replicate a previous study, and spatial regression models are employed to explore discontinuities at borders. Results: Cancer rate differences lack statistical significance within time zones and near borders for total cancer and most individual cancers. Exceptions included breast, prostate, and liver \& bile duct cancers, which exhibited significant relationships with relative position at the 95\% significance level. Breast and liver and bile duct cancers saw decreases, while prostate cancer incidence increased from west to east within time zones. Conclusions: Relative position does not have a significant impact on cancer incidence, hence cancer development in general. Isolated exceptions may warrant further investigation as more data becomes available. Impact: Our findings challenge prior research, revealing numerous inconsistencies. These disparities urge a reconsideration of the potential disparities in human health associated with daylight saving time and standard time. They offer insights contribute to the ongoing discussion surrounding the retention or abandonment of DST.

stat.AP↗

Generalized maximum likelihood estimation of the mean of parameters of mixtures, with applications to sampling

Let $f(y|θ), \; θ\in Ω$ be a parametric family, $η(θ)$ a given function, and $G$ an unknown mixing distribution. It is desired to estimate $E_G (η(θ))\equiv η_G$ based on independent observations $Y_1,...,Y_n$, where $Y_i \sim f(y|θ_i)$, and $θ_i \sim G$ are iid. We explore the Generalized Maximum Likelihood Estimators (GMLE) for this problem. Some basic properties and representations of those estimators are shown. In particular we suggest a new perspective, of the weak convergence result by Kiefer and Wolfowitz (1956), with implications to a corresponding setup in which $θ_1,...,θ_n$ are {\it fixed} parameters. We also relate the above problem, of estimating $η_G$, to non-parametric empirical Bayes estimation under a squared loss. Applications of GMLE to sampling problems are presented. The performance of the GMLE is demonstrated both in simulations and through a real data example.

math.ST↗

Rank-Constrained Least-Squares: Prediction and Inference

In this work, we focus on the high-dimensional trace regression model with a low-rank coefficient matrix. We establish a nearly optimal in-sample prediction risk bound for the rank-constrained least-squares estimator under no assumptions on the design matrix. Lying at the heart of the proof is a covering number bound for the family of projection operators corresponding to the subspaces spanned by the design. By leveraging this complexity result, we perform a power analysis for a permutation test on the existence of a low-rank signal under the high-dimensional trace regression model. We show that the permutation test based on the rank-constrained least-squares estimator achieves non-trivial power with no assumptions on the minimum (restricted) eigenvalue of the covariance matrix of the design. Finally, we use alternating minimization to approximately solve the rank-constrained least-squares problem to evaluate its empirical in-sample prediction risk and power of the resulting permutation test in our numerical study.

math.ST↗

Estimation of a score-explained non-randomized treatment effect in fixed and high dimensions

Non-randomized treatment effect models are widely used for the assessment of treatment effects in various fields and in particular social science disciplines like political science, psychometry, psychology. More specifically, these are situations where treatment is assigned to an individual based on some of their characteristics (e.g. scholarship is allocated based on merit or antihypertensive treatments are allocated based on blood pressure level) instead of being allocated randomly, as is the case, for example, in randomized clinical trials. Popular methods that have been largely employed till date for estimation of such treatment effects suffer from slow rates of convergence (i.e. slower than $\sqrt{n}$). In this paper, we present a new model coined SCENTS: Score Explained Non-Randomized Treatment Systems, and a corresponding method that allows estimation of the treatment effect at $\sqrt{n}$ rate in the presence of fairly general forms of confoundedness, when the `score' variable on whose basis treatment is assigned can be explained via certain feature measurements of the individuals under study. We show that our estimator is asymptotically normal in general and semi-parametrically efficient under normal errors. We further extend our analysis to high dimensional covariates and propose a $\sqrt n$ consistent and asymptotically normal estimator based on a de-biasing procedure. Our analysis for the high dimensional incarnation can be readily extended to analyze partial linear models in the presence of noisy variables corresponding to the non-linear part of the model, where the noise can be correlated with the variables corresponding to the linear part. We analyze two real datasets via our method and compare our results with those obtained by using previous approaches. We conclude this paper with a discussion on some possible extensions of our approach.

stat.ME↗

Nonparametric Empirical Bayes Estimation and Testing for Sparse and Heteroscedastic Signals

Large-scale modern data often involves estimation and testing for high-dimensional unknown parameters. It is desirable to identify the sparse signals, ``the needles in the haystack'', with accuracy and false discovery control. However, the unprecedented complexity and heterogeneity in modern data structure require new machine learning tools to effectively exploit commonalities and to robustly adjust for both sparsity and heterogeneity. In addition, estimates for high-dimensional parameters often lack uncertainty quantification. In this paper, we propose a novel Spike-and-Nonparametric mixture prior (SNP) -- a spike to promote the sparsity and a nonparametric structure to capture signals. In contrast to the state-of-the-art methods, the proposed methods solve the estimation and testing problem at once with several merits: 1) an accurate sparsity estimation; 2) point estimates with shrinkage/soft-thresholding property; 3) credible intervals for uncertainty quantification; 4) an optimal multiple testing procedure that controls false discovery rate. Our method exhibits promising empirical performance on both simulated data and a gene expression case study.

cs.LG↗

High-Dimensional Varying Coefficient Models with Functional Random Effects

We consider a sparse high-dimensional varying coefficients model with random effects, a flexible linear model allowing covariates and coefficients to have a functional dependence with time. For each individual, we observe discretely sampled responses and covariates as a function of time as well as time invariant covariates. Under sampling times that are either fixed and common or random and independent amongst individuals, we propose a projection procedure for the empirical estimation of all varying coefficients. We extend this estimator to construct confidence bands for a fixed number of varying coefficients.

math.ST↗

Asymptotic normality of a linear threshold estimator in fixed dimension with near-optimal rate

Linear thresholding models postulate that the conditional distribution of a response variable in terms of covariates differs on the two sides of a (typically unknown) hyperplane in the covariate space. A key goal in such models is to learn about this separating hyperplane. Exact likelihood or least squares methods to estimate the thresholding parameter involve an indicator function which make them difficult to optimize and are, therefore, often tackled by using a surrogate loss that uses a smooth approximation to the indicator. In this paper, we demonstrate that the resulting estimator is asymptotically normal with a near optimal rate of convergence: $n^{-1}$ up to a log factor, in both classification and regression thresholding models. This is substantially faster than the currently established convergence rates of smoothed estimators for similar models in the statistics and econometrics literatures. We also present a real-data application of our approach to an environmental data set where $CO_2$ emission is explained in terms of a separating hyperplane defined through per-capita GDP and urban agglomeration.

math.ST↗