SearcharxivSearch

arXiv subjects

Youheng Zhu

Publications and source records attributed to Youheng Zhu.

5 recordsLinked to original sources

High-Dimensional Interpolators Can Be Fragile: Heavy Tails and High-Dimensional Large Deviations

High-dimensional interpolation is common in modern machine learning, but its tail risk is less understood than its expected prediction risk. Existing theory shows that interpolating models can perform well in expectation, yet such guarantees do not determine the probability of rare, severe errors. In operations research and stochastic decision-making applications, rare estimation errors can have disproportionate downstream effects, so tail behavior matters alongside average performance. We study the fragility of high-dimensional linear interpolators using large-deviation methods. We focus on ridgeless regression and compare it with ridge-regularized estimators. We first show that the risk of ridgeless regression can exhibit heavy-tailed behavior: although its expected risk may remain well controlled, its upper tail can decay much more slowly than that of regularized alternatives. We then quantify this phenomenon at the level of large-deviation rates. In the regime we study, ridge regularization suppresses fixed right-tail deviations at the $n^2$ scale, whereas ridgeless regression has only $n\log n$-scale decay, where $n$ is the sample size. This gap shows that interpolation can be statistically fragile even when it is accurate on average. Thus regularization affects the frequency of rare, high-impact risk events in addition to the usual bias-variance tradeoff.

math.ST

An AI-Assisted Solution to the Signed BAR Conjecture: Uniqueness in the Harrison--Reiman Class and a Completely-$\mathcal{S}$ Class Obstruction

For a multidimensional reflected diffusion, determining whether the associated basic adjoint relationship (BAR) uniquely characterizes the stationary distribution is a basic uniqueness problem in the BAR approach. The problem has remained unresolved for more than 35 years since the introduction of the BAR approach. In this paper, we resolve the finite-signed uniqueness problem for stable Harrison--Reiman data with a nonsingular $M$-matrix reflection matrix. The proof uses pathwise differentiability of the reflected diffusion implies feasible directional differentiability of the probabilistic resolvent to show that, at boundary points, its one-sided initial-state derivative factors through the tangent projection and vanishes along active reflection directions. An interior one-sided convolution then yields smooth test functions whose oblique derivatives are uniformly bounded and converge pointwise to zero on each closed face. The interior signed measure is consequently invariant for the reflected semigroup. The proof was discovered with the assistance of ChatGPT 5.5 Pro and subsequently verified by the authors. We also show that the nonsingular $M$-matrix assumption is structural. In the larger completely-$\mathcal{S}$ class, a nonsingular reflection matrix with a singular proper principal block admits boundary gauges supported on lower-dimensional strata. Under standard exponential ergodicity and a mild one-step regulator bound, these gauges produce nonzero zero-mass signed BAR tuples; indeed the zero-mass interior BAR coordinates contain an infinite-dimensional subspace. A four-parameter three-dimensional family, including an explicit rational example, verifies the obstruction. Thus the finite signed version of the Dai--Dieker question has a positive answer in the Harrison--Reiman $M$-matrix class and a negative answer in a natural completely-$\mathcal{S}$ extension.

math.PR

A Covering Framework for Offline POMDPs Learning using Belief Space Metric

In off policy evaluation (OPE) for partially observable Markov decision processes (POMDPs), an agent must infer hidden states from past observations, which exacerbates both the curse of horizon and the curse of memory in existing OPE methods. This paper introduces a novel covering analysis framework that exploits the intrinsic metric structure of the belief space (distributions over latent states) to relax traditional coverage assumptions. By assuming value relevant functions are Lipschitz continuous in the belief space, we derive error bounds that mitigate exponential blow ups in horizon and memory length. Our unified analysis technique applies to a broad class of OPE algorithms, yielding concrete error bounds and coverage requirements expressed in terms of belief space metrics rather than raw history coverage. We illustrate the improved sample efficiency of this framework via case studies: the double sampling Bellman error minimization algorithm, and the memory based future dependent value functions (FDVF). In both cases, our coverage definition based on the belief space metric yields tighter bounds.

stat.ML

On the Power of (Approximate) Reward Models for Inference-Time Scaling

Inference-time scaling has recently emerged as a powerful paradigm for improving the reasoning capability of large language models. Among various approaches, Sequential Monte Carlo (SMC) has become a particularly important framework, enabling iterative generation, evaluation, rejection, and resampling of intermediate reasoning trajectories. A central component in this process is the reward model, which evaluates partial solutions and guides the allocation of computation during inference. However, in practice, true reward models are never available. All deployed systems rely on approximate reward models, raising a fundamental question: Why and when do approximate reward models suffice for effective inference-time scaling? In this work, we provide a theoretical answer. We identify the Bellman error of the approximate reward model as the key quantity governing the effectiveness of SMC-based inference-time scaling. For a reasoning process of length $T$, we show that if the Bellman error of the approximate reward model is bounded by $O(1/T)$, then combining this reward model with SMC reduces the computational complexity of reasoning from exponential in $T$ to polynomial in $T$. This yields an exponential improvement in inference efficiency despite using only approximate rewards.

cs.CL

Information-theoretic Analysis of the Gibbs Algorithm: An Individual Sample Approach

Recent progress has shown that the generalization error of the Gibbs algorithm can be exactly characterized using the symmetrized KL information between the learned hypothesis and the entire training dataset. However, evaluating such a characterization is cumbersome, as it involves a high-dimensional information measure. In this paper, we address this issue by considering individual sample information measures within the Gibbs algorithm. Our main contribution lies in establishing the asymptotic equivalence between the sum of symmetrized KL information between the output hypothesis and individual samples and that between the hypothesis and the entire dataset. We prove this by providing explicit expressions for the gap between these measures in the non-asymptotic regime. Additionally, we characterize the asymptotic behavior of various information measures in the context of the Gibbs algorithm, leading to tighter generalization error bounds. An illustrative example is provided to verify our theoretical results, demonstrating our analysis holds in broader settings.

cs.IT