SearcharxivSearch

arXiv subjects

Yanjun Han

Publications and source records attributed to Yanjun Han.

At least 19 recordsLinked to original sources

A Cap-Move Reformulation of Ling's Proof of Samuels' Conjecture

Let $0\le \mu_1\le \mu_2 \le \cdots \le \mu_n$ and $\delta > 0$. Samuels' conjecture claims that if $X_1,\dots,X_n$ are independent non-negative random variables with $\mathbb{E}[X_i] = \mu_i$, then $$ \mathbb{P}\left( \sum_{i=1}^n X_i < \delta + \sum_{i=1}^n \mu_i \right) \ge \min_{1\le i\le n} \prod_{j=i}^n \left(1-\frac{\mu_j}{\delta + \sum_{k=i}^n \mu_k}\right).$$ This conjecture was recently proved by Ling. In this note, we provide an alternative presentation of Ling's proof via an operation called the \emph{cap move}. This proof works directly with finitely supported distributions and avoids the reduction to the Bernoulli case.

math.PR

Sharp small-deviation inequalities for sums of independent nonnegative random variables

Let $(X_1,\ldots,X_n)$ be independent nonnegative random variables with $\mathbb{E} X_i\le1$, and write $S=\sum_iX_i$. For $\delta>0$, we prove that \[ \mathbb{P}\left(S<\mathbb{E} S+\delta\right)\ge b_{n,\delta}, \] where $b_{n,\delta}=\delta(n/(n+\delta))^n$ for $0<\delta<1$ and $b_{n,\delta}=(1-1/(n+\delta))^n$ for $\delta\ge1$. The bound is sharp for every $n$ and $\delta\ge 1$. In particular, since $b_{n,\delta} \ge e^{-1}$ for $\delta \ge 1$, our result proves Feige's conjecture [Feige, 2004] in the affirmative for $\delta\ge 1$. The proof is found by ChatGPT 5.6 Pro. It combines the exact Dirichlet calibration theorem of Vlassis and Thomas [Vlassis and Thomas, 2026], which resolves Gaffke's conjecture in statistics, with results in convex geometry including Gr\"unbaum's centroid theorem [Gr\"unbaum, 1960] and its generalization by Letwin and Yaskin [Letwin and Yaskin, 2024].

math.PR

Elementary Symmetric Polynomial Inequalities for Centered Vectors and Matrices

We prove new inequalities for elementary symmetric polynomials (ESPs) for vectors that sum to zero, and for square matrices with zero row and column sums. We apply these results to obtain a unified upper bound on the mean-field approximation guarantee for permutation mixtures, as well as a sharp $\chi^2$ version of the de Finetti theorem for finite sequences over a small alphabet. The main proof ideas were developed by the GPT-5.5 Pro model.

math.CO

The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex Optimization

We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for $1$-Lipschitz functions on the $d$-dimensional Euclidean ball. For time horizons $n\ge d^{10/3}$, we prove a lower bound of $\Omega(d^{4/3}\sqrt{n})$, the first nontrivial bound that exceeds the $d\sqrt{n}$ dependence of linear bandits, showing that stochastic bandit convex optimization is fundamentally harder than linear bandits. For $d^2\le n\le d^{10/3}$, we obtain a lower bound of $\Omega(\sqrt{d}n^{3/4})$, matching the regret of the algorithm of Flaxman et al. (2005), establishing its optimality in this regime. The hard class of convex functions we construct takes the following form in dimension $2d$: for an action $a=(a^1,a^2)\in \mathbb{B}^{2d}$, each function is the scaled soft maximum of a "tube", $r^{-1}\|W^\star a^1-\frac{r}{8\varepsilon}a^2 \|$ (hyperparameterized by $\varepsilon,r$), and a squared distance function, $\frac12\|a^1-u^\star\|^2-\frac12\|u^\star\|^2$. Here $u^\star\in\mathbb{R}^d$ is the unknown target determining the minimizer, while $W^\star\in\mathbb{R}^{d\times d}$ hides the region in which the quadratic curvature is observable. Indeed, observations reveal substantial information about $u^\star$ only when the learner acts near the hidden tube $a^2\approx \frac{8\varepsilon}{r}W^\star a^1$; away from it, the tube branch masks the quadratic branch. Thus the learner must pay to uncover the geometry encoded by $W^\star$ before it can effectively exploit the curvature that identifies $u^\star$. Formalizing this tradeoff yields a sample complexity lower bound of $\Omega(\frac{d^{5/2}}{\varepsilon^2}\wedge\frac{d^2}{\varepsilon^4})$ for finding an $\varepsilon$-optimal action, and ultimately the $\Omega(d^{4/3}\sqrt{n}\wedge\sqrt{d}n^{3/4})$ regret lower bound. The proof was developed by GPT-5.5 Pro and GPT-5.6 Sol Pro under the authors' guidance.

stat.ML

The (Marginal) Value of a Search Ad: An Online Causal Framework for Repeated Second-price Auctions

Existing auto-bidding algorithms in digital advertising often treat the value of an ad opportunity as the revenue obtained when an ad is shown and/or clicked, and bid accordingly. This can lead to wasteful spending because the true value is the marginal gain from paid exposure: even without winning a sponsored slot, an advertiser may still earn revenue via an organic search result (e.g., on Google or Amazon). Motivated by recent work, we model ad value as a treatment effect--the outcome difference between winning and losing the auction--and study online learning for bidding in second-price (Vickrey) auctions under this causal perspective. We develop algorithms that attain rate-optimal regret under several feedback models. A key ingredient exploits the information revealed by the second-price payment rule, which strictly improves regret relative to analogous learning problems in first-price auctions.

cs.GT

An Empirical Bayes Perspective on Heteroskedastic Mean Estimation

Towards understanding the fundamental limits of estimation from data of varied quality, we study the problem of estimating a mean parameter from heteroskedastic Gaussian observations where the variances are unknown and may vary arbitrarily across observations. While a simple linear estimator with known variances attains the smallest mean squared error, estimation without this knowledge is challenging due to the large number of nuisance parameters. We propose a simple and principled approach based on empirical Bayes: model the observations as if they were i.i.d. from a normal scale mixture and compute the profile maximum likelihood estimator (MLE) for the mean, treating the nonparametric mixing distribution as nuisance. Our result shows that this estimator achieves near-optimal error bounds across various heteroskedastic models in the literature. In particular, for the subset-of-signals problem where an unknown subset of observations has small variance, our estimator adaptively achieves the minimax rate for all signal sizes, including the sharp phase transition, without any tuning parameters. One of our key technical steps is a sharper metric entropy bound for normal scale mixtures, obtained via Chebyshev approximations on a transformed polynomial basis. This approach yields an improved polylogarithmic, rather than polynomial, dependence on the variance ratio, which could be of independent interest.

math.ST

Interactive Learning of Single-Index Models via Stochastic Gradient Descent

Stochastic gradient descent (SGD) is a cornerstone algorithm for high-dimensional optimization, renowned for its empirical successes. Recent theoretical advances have provided a deep understanding of how SGD enables feature learning in high-dimensional nonlinear models, most notably the \textit{single-index model} with i.i.d. data. In this work, we study the sequential learning problem for single-index models, also known as generalized linear bandits or ridge bandits, where SGD is a simple and natural solution, yet its learning dynamics remain largely unexplored. We show that, similar to the optimal interactive learner, SGD undergoes a distinct ``burn-in'' phase before entering the ``learning'' phase in this setting. Moreover, with an appropriately chosen learning rate schedule, a single SGD procedure simultaneously achieves near-optimal (or best-known) sample complexity and regret guarantees across both phases, for a broad class of link functions. Our results demonstrate that SGD remains highly competitive for learning single-index models under adaptive data.

stat.ML

PETS: A Principled Framework Towards Optimal Trajectory Allocation for Efficient Test-Time Self-Consistency

Test-time scaling can improve model performance by aggregating stochastic reasoning trajectories. However, achieving sample-efficient test-time self-consistency under a limited budget remains an open challenge. We introduce PETS (Principled and Efficient Test-TimeSelf-Consistency), which initiates a principled study of trajectory allocation through an optimization framework. Central to our approach is the self-consistency rate, a new measure defined as agreement with the infinite-budget majority vote. This formulation makes sample-efficient test-time allocation theoretically grounded and amenable to rigorous analysis. We study both offline and online settings. In the offline regime, where all questions are known in advance, we connect trajectory allocation to crowdsourcing, a classic and well-developed area, by modeling reasoning traces as workers. This perspective allows us to leverage rich existing theory, yielding theoretical guarantees and an efficient majority-voting-based allocation algorithm. In the online streaming regime, where questions arrive sequentially and allocations must be made on the fly, we propose a novel method inspired by the offline framework. Our approach adapts budgets to question difficulty while preserving strong theoretical guarantees and computational efficiency. Experiments show that PETS consistently outperforms uniform allocation. On GPQA, PETS achieves perfect self-consistency in both settings while reducing the sampling budget by up to 75% (offline) and 55% (online) relative to uniform allocation. Code is available at https://github.com/ZDCSlab/PETS.

cs.LG

Universal priors: solving empirical Bayes via Bayesian inference and pretraining

We theoretically justify the recent empirical finding of [Teh et al., 2025] that a transformer pretrained on synthetically generated data achieves strong performance on empirical Bayes (EB) problems. We take an indirect approach to this question: rather than analyzing the model architecture or training dynamics, we ask why a pretrained Bayes estimator, trained under a prespecified training distribution, can adapt to arbitrary test distributions. Focusing on Poisson EB problems, we identify the existence of universal priors such that training under these priors yields a near-optimal regret bound of $\widetilde{O}(\frac{1}{n})$ uniformly over all test distributions. Our analysis leverages the classical phenomenon of posterior contraction in Bayesian statistics, showing that the pretrained transformer adapts to unknown test distributions precisely through posterior contraction. This perspective also explains the phenomenon of length generalization, in which the test sequence length exceeds the training length, as the model performs Bayesian inference using a generalized posterior.

stat.ML

Optimal Arm Elimination Algorithms for Combinatorial Bandits

Combinatorial bandits extend the classical bandit framework to settings where the learner selects multiple arms in each round, motivated by applications such as online recommendation and assortment optimization. While extensions of upper confidence bound (UCB) algorithms arise naturally in this context, adapting arm elimination methods has proved more challenging. We introduce a novel elimination scheme that partitions arms into three categories (confirmed, active, and eliminated), and incorporates explicit exploration to update these sets. We demonstrate the efficacy of our algorithm in two settings: the combinatorial multi-armed bandit with general graph feedback, and the combinatorial linear contextual bandit. In both cases, our approach achieves near-optimal regret, whereas UCB-based methods can provably fail due to insufficient explicit exploration. Matching lower bounds are also provided.

cs.LG

Sharp mean-field analysis of permutation mixtures and permutation-invariant decisions

We develop sharp bounds on the statistical distance between high-dimensional permutation mixtures and their i.i.d. counterparts. Our approach establishes a new geometric link between the spectrum of a complex channel overlap matrix and the information geometry of the channel, yielding tight dimension-independent bounds that close gaps left by previous work. Within this geometric framework, we also derive dimension-dependent bounds that uncover phase transitions in dimensionality for Gaussian and Poisson families. Applied to compound decision problems, this refined control of permutation mixtures enables sharper mean-field analyses of permutation-invariant decision rules, yielding strong non-asymptotic equivalence results between two notions of compound regret in Gaussian and Poisson models.

math.ST

Besting Good--Turing: Optimality of Non-Parametric Maximum Likelihood for Distribution Estimation

When faced with a small sample from a large universe of possible outcomes, scientists often turn to the venerable Good--Turing estimator. Despite its pedigree, however, this estimator comes with considerable drawbacks, such as the need to hand-tune smoothing parameters and the lack of a precise optimality guarantee. We introduce a parameter-free estimator that bests Good--Turing in both theory and practice. Our method marries two classic ideas, namely Robbins's empirical Bayes and Kiefer--Wolfowitz non-parametric maximum likelihood estimation (NPMLE), to learn an implicit prior from data and then convert it into probability estimates. We prove that the resulting estimator attains the optimal instance-wise risk up to logarithmic factors in the competitive framework of Orlitsky and Suresh, and that the Good--Turing estimator is strictly suboptimal in the same framework. Our simulations on synthetic data and experiments with English corpora and U.S. Census data show that our estimator consistently outperforms both the Good--Turing estimator and explicit Bayes procedures.

math.ST

Evolution of Information in Interactive Decision Making: A Case Study for Multi-Armed Bandits

We study the evolution of information in interactive decision making through the lens of a stochastic multi-armed bandit problem. Focusing on a fundamental example where a unique optimal arm outperforms the rest by a fixed margin, we characterize the optimal success probability and mutual information over time. Our findings reveal distinct growth phases in mutual information -- initially linear, transitioning to quadratic, and finally returning to linear -- highlighting curious behavioral differences between interactive and non-interactive environments. In particular, we show that optimal success probability and mutual information can be decoupled, where achieving optimal learning does not necessarily require maximizing information gain. These findings shed new light on the intricate interplay between information and learning in interactive decision making.

stat.ML

Joint Value Estimation and Bidding in Repeated First-Price Auctions

We study regret minimization in repeated first-price auctions (FPAs), where a bidder observes only the realized outcome after each auction -- win or loss. This setup reflects practical scenarios in online display advertising where the actual value of an impression depends on the difference between two potential outcomes, such as clicks or conversion rates, when the auction is won versus lost. We incorporate causal inference into this framework and analyze the challenging case where only the treatment effect admits a simple dependence on observable features. Under this framework, we propose algorithms that jointly estimate private values and optimize bidding strategies under two different feedback types on the highest other bid (HOB): the full-information feedback where the HOB is always revealed, and the binary feedback where the bidder only observes the win-loss indicator. Under both cases, our algorithms are shown to achieve near-optimal regret bounds. Notably, our framework enjoys a unique feature that the treatments are actively chosen, and hence eliminates the need for the overlap condition commonly required in causal inference.

cs.LG

Reconfigurable nonlinear optical computing device for retina-inspired computing

Optical neural networks are at the forefront of computational innovation, utilizing photons as the primary carriers of information and employing optical components for computation. However, the fundamental nonlinear optical device in the neural networks is barely satisfied because of its high energy threshold and poor reconfigurability. This paper proposes and demonstrates an optical sigmoid-type nonlinear computation mode of Vertical-Cavity Surface-Emitting Lasers (VCSELs) biased beneath the threshold. The device is programmable by simply adjusting the injection current. The device exhibits sigmoid-type nonlinear performance at a low input optical power ranging from merely 3-250 {\mu}W. The tuning sensitivity of the device to the programming current density can be as large as 15 {\mu}W*mm2/mA. Deep neural network architecture based on such device has been proposed and demonstrated by simulation on recognizing hand-writing digital dataset, and a 97.3% accuracy has been achieved. A step further, the nonlinear reconfigurability is found to be highly useful to enhance the adaptability of the networks, which is demonstrated by significantly improving the recognition accuracy by 41.76%, 19.2%, and 25.89% of low-contrast hand-writing digital images under high exposure, low exposure, and high random noise respectively.

physics.optics

D-band MUTC Photodiode Module for Ultra-Wideband 160 Gbps Photonics-Assisted Fiber-THz Integrated Communication System

Current wireless communication systems are increasingly constrained by insufficient bandwidth and limited power output, impeding the achievement of ultra-high-speed data transmission. The terahertz (THz) range offers greater bandwidth, but it also imposes higher requirements on broadband and high-power devices. In this work, we present a modified uni-traveling-carrier photodiode (MUTC-PD) module with WR-6 waveguide output for photonics-assisted fiber-THz integrated wireless communications. Through the optimization of the epitaxial structure and high-impedance coplanar waveguide (CPW), the fabricated 6-um-diameter MUTC-PD achieves a high output power of -0.96 dBm at 150 GHz and ultra-flat frequency response at D-band. The MUTC-PD is subsequently packaged into a compact WR-6 module, incorporating planar-circuit-based RF-choke, DC-block and probe. The packaged PD module demonstrates high saturation power and flat frequency responses with minimal power roll-off of only 2 dB over 110-170 GHz. By incorporating the PD module into a fiber-THz integrated communication system, high data rates of up to 160 Gbps with 16 quadrature amplitude modulation (QAM) and a maximum symbol transmission rate of 60 Gbaud with QPSK modulation are successfully secured. The demonstration verifies the potential of the PD module for ultra-broadband and ultra-high-speed THz communications, setting a foundation for future research in high-speed data transmission.

physics.optics

Ultra-High-Efficiency Dual-Band Thin-Film Lithium Niobate Modulator Incorporating Low-k Underfill with 220 GHz Extrapolated Bandwidth for 390 Gbit/s PAM8 Transmission

High-performance electro-optic modulators play a critical role in modern telecommunication networks and intra-datacenter interconnects. Low driving voltage, large electro-optic bandwidth, compact device size, and multi-band operation ability are essential for various application scenarios, especially energy-efficient high-speed data transmission. However, it is challenging to meet all these requirements simultaneously. Here, we demonstrate a high-performance dual-band thin-film lithium niobate electro-optic modulator with low-k underfill to achieve overall performance improvement. The low-k material helps reduce the RF loss of the modulator and achieve perfect velocity matching with narrow electrode gap to overcome the voltage-bandwidth limitation, extending electro-optic bandwidth and enhancing modulation efficiency simultaneously. The fabricated 7-mm-long modulator exhibits a low half-wave voltage of 1.9 V at C-band and 1.54 V at O-band, featuring a low half-wave voltage-length product of 1.33 V*cm and 1.08 V*cm, respectively. Meanwhile, the novel design yields an ultra-wide extrapolated 3 dB bandwidth of 220 GHz (218 GHz) in the C-band (O-band). High-speed data transmission in both C- and O-bands using the same device has been demonstrated for the first time by PAM8 with data rates up to 390 Gbit/s, corresponding to a record-low energy consumption of 0.69 fJ/bit for next-generation cost-effective ultra-high-speed optical communications.

physics.optics

Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability

We develop a unifying framework for information-theoretic lower bound in statistical estimation and interactive decision making. Classical lower bound techniques -- such as Fano's method, Le Cam's method, and Assouad's lemma -- are central to the study of minimax risk in statistical estimation, yet are insufficient to provide tight lower bounds for \emph{interactive decision making} algorithms that collect data interactively (e.g., algorithms for bandits and reinforcement learning). Recent work of Foster et al. (2021, 2023) provides minimax lower bounds for interactive decision making using seemingly different analysis techniques from the classical methods. These results -- which are proven using a complexity measure known as the \emph{Decision-Estimation Coefficient} (DEC) -- capture difficulties unique to interactive learning, yet do not recover the tightest known lower bounds for passive estimation. We propose a unified view of these distinct methodologies through a new lower bound approach called \emph{interactive Fano method}. As an application, we introduce a novel complexity measure, the \emph{Fractional Covering Number}, which facilitates the new lower bounds for interactive decision making that extend the DEC methodology by incorporating the complexity of estimation. Using the fractional covering number, we (i) provide a unified characterization of learnability for \emph{any} stochastic bandit problem, (ii) close the remaining gap between the upper and lower bounds in Foster et al. (2021, 2023) (up to polynomial factors) for any interactive decision making problem in which the underlying model class is convex.

cs.LG