SearcharxivSearch

arXiv subjects

Yu-Jie Zhang

Publications and source records attributed to Yu-Jie Zhang.

At least 19 recordsLinked to original sources

Efficient Multinomial Logistic Bandit via Frequent Directions

This paper studies efficient online algorithms for multinomial logistic bandits (MLogB), where the feedback distribution over $K+1$ outcomes follows a multinomial logistic model of $d$-dimensional action vectors. A representative UCB-type algorithm, OFUL-MLogB, achieves a regret bound of $\tilde{\mathcal{O}}(Kd\sqrt{T})$, but still requires $\mathcal{O}(K^3d^3)$ time and $\mathcal{O}(K^2d^2)$ space per round due to parameter estimation and optimistic reward construction, which is prohibitive in high-dimensional settings. To address this limitation, we propose EOFD-MLogB, which integrates frequent directions matrix sketching into OFUL-MLogB. By maintaining a low-rank SVD sketch of the accumulated Hessian, constrained online Newton updates in parameter estimation and $Kd \times K$ spectral-norm computations in the reward bonus are reduced to one-dimensional root-finding tasks and $K \times K$ eigenvalue computations, respectively. This yields dominant per-round time complexity $\mathcal{O}(Kd(m+K)^2)$ and space complexity $\mathcal{O}(Kd(m+K))$, where $m \ll d$ is the sketch size. We further prove a regret bound of $\tilde{\mathcal{O}}(Δ_T(Kd\lnΔ_T+m)\sqrt{T})$, where the sketching error factor $Δ_T$ is controlled by the $m$-truncated spectral tail of the Hessian. Thus, when the Hessian is approximately low-rank, the regret is close to that of OFUL-MLogB. Experiments validate the computational efficiency and competitive performance.

cs.LG

Near-Optimal Regret in Adversarial Kernel Bandits

We study the adversarial kernel bandit problem, in which the loss at each round is induced by an arbitrary bounded element of a reproducing kernel Hilbert space (RKHS). We propose an exponential-weights algorithm built on a regularized importance-weighted loss estimator, together with an explicit correction term that cancels the bias introduced by the regularization. Our main result bounds the regret by $\widetilde{O}\big(\sqrt{T\, d_*(λ)\,\log|{X}|}\big)$, where $d_*(λ)$ is a widely-adopted notion of effective dimension that captures the complexity of the kernel. Up to logarithmic factors, this matches the known rate achieved in the related stochastic kernel bandit problem. A notable application is the Matérn$(ν,d)$ kernel with smoothness parameter $ν$ on $\mathbb{R}^d$, for which our bound specializes to $\widetilde{O}\big(T^{(ν+d)/(2ν+d)}\big)$, improving over the best-known prior rate of Chatterji et al. [2019] while simultaneously removing the rank-one adversary assumption required by their analysis. Moreover, this rate is the same as the known optimal rate for stochastic kernel bandits, and also matches a lower bound from concurrent work up to a $\log T$ factor.

cs.LG

Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer

We study dynamic regret minimization in non-stationary online learning, with a primary focus on follow-the-regularized-leader (FTRL) methods. FTRL is important for curved losses and for understanding adaptive optimizers such as Adam, yet existing dynamic regret analyses are less explored for FTRL. To address this, we build on the discounted-to-dynamic reduction and present a modular way to obtain dynamic regret bounds of FTRL-related problems. Specifically, we focus on two representative curved losses: linear regression and logistic regression. Our method not only simplifies existing proofs for the optimal dynamic regret of online linear regression, but also yields new dynamic regret guarantees for online logistic regression. Beyond online convex optimization, we apply the reduction to analyze the Adam optimizers, obtaining optimal convergence rates in stochastic, non-convex, and non-smooth settings. The reduction also enables a more detailed treatment of Adam with two discount parameters $(β_1,β_2)$, leading to new results for both clipped and clip-free variants of Adam optimizers.

cs.LG

Probing the $γγ^*\to η^{(\prime)}$ Transition Form Factors with Newly Derived $η^{(\prime)}$-Meson Light-Cone Distribution Amplitudes

In the present work, we analyze the properties of the transition form factors (TFFs) for the $γγ^*\to η^{(\prime)}$ process, employing the $η^{(\prime)}$-meson light-cone distribution amplitude (LCDA) derived within the light-cone sum rule framework. To this end, we adopt the quark-flavor mixing scheme for the $η^{(\prime)}$ meson, and compute the TFFs by systematically incorporating transverse-momentum corrections and contributions beyond the leading Fock state. We utilize light-cone harmonic oscillator models to parameterize the longitudinal and transverse behavior of the leading-twist light-cone wavefunction, for which the corresponding LCDA exhibits a unimodal profile. We further examine the potential contributions of intrinsic charm components to the scaled TFFs $Q^2 F_{ηγ}(Q^2)$ and $Q^2 F_{η^\prime γ}(Q^2)$. Leveraging a range of values for the decay constant $f_{η_{c_0}}$ and implementing the $η$-$η'$-$η_c$ and $η$-$η^\prime$-$G$-$η_c$ mixing mechanisms accordingly, together with the recently updated mixing angles, we investigate the impact of the intrinsic $c\bar{c}$ and gluonic component on these observables. In high-$Q^2$ regime, $Q^2 F_{η^\primeγ}(Q^2)$ exhibits a marked increase in sensitivity to the charm quark component, whereas $Q^2F_{ηγ}(Q^2)$ becomes notably stabilized. A detailed discussion of $χ^2/d.o.f$ and $p$-values indicates that the intrinsic charm quark component is important and yields a substantial, non-negligible contribution across the entire $Q^2$ range.

hep-ph

Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update

We study the generalized linear bandit (GLB) problem, a contextual multi-armed bandit framework that extends the classical linear model by incorporating a non-linear link function, thereby modeling a broad class of reward distributions such as Bernoulli and Poisson. While GLBs are widely applicable to real-world scenarios, their non-linear nature introduces significant challenges in achieving both computational and statistical efficiency. Existing methods typically trade off between two objectives, either incurring high per-round costs for optimal regret guarantees or compromising statistical efficiency to enable constant-time updates. In this paper, we propose a jointly efficient algorithm that attains a nearly optimal regret bound with $\mathcal{O}(1)$ time and space complexities per round. The core of our method is a tight confidence set for the online mirror descent (OMD) estimator, which is derived through a novel analysis that leverages the notion of mix loss from online prediction. The analysis shows that our OMD estimator, even with its one-pass updates, achieves statistical efficiency comparable to maximum likelihood estimation, thereby leading to a jointly efficient optimistic method.

cs.LG

$η$-$η'$ mixing and its application in the $B^+/D^+/D_s^+\toη^{(\prime)}\ell^+ ν_\ell$ decays

In this paper, we take into account the intrinsic charm and gluonic contents into the $η-η^\prime$ mixing scheme and formulate the tetramixing $η-η^\prime-G-η_c$ to study the mixing properties of $η^{(\prime)}$ mesons. Using the newly derived mixing parameters, we calculate the transition form factors (TFFs) of $B^+/D^+/D_s^+\toη^{(\prime)}$ within the QCD light-cone sum rules up to next-to-leading order QCD corrections and twist-4 contributions. Using the extrapolated TFFs, we then calculate the decay widths and branching fractions of the semi-leptonic decays $B^+/D^+/D_s^+\toη^{(\prime)}\ell^+ν_{\ell}$. Our results are consistent with the recent Belle and BES-III measurements within reasonable errors.

hep-ph

Triple top baryon $Ω_{ttt}$

The recent observation of toponium by CMS and ATLAS has renewed interest in top quark bound states. In this work, we present an exploratory but quantitative study of the hypothetical triple-top baryon, denoted as $Ω_{ttt}$, the only baryon that is governed by ultraviolet freedom. Using a variational method with an effective potential of $ttt$ that includes QCD, Higgs, and QED contributions, we estimate its mass to be around 514 GeV with a binding energy of about 4 GeV. We further discuss its possible production at future high-energy colliders, finding that the cross sections are extremely suppressed. The dominant weak decay channel is identified as $Ω_{ttt}\to W^+W^+W^+bbb$, leading to complex multi-lepton and multi-jet final states. Our analysis, though approximate, demonstrates the distinctive features of $Ω_{ttt}$ compared with other triply-heavy baryons such as $Ω_{ccc}$ and $Ω_{bbb}$, and may serve as a starting point for more refined approaches, including lattice QCD or effective field theory. This work highlights both the theoretical challenges and the potential opportunities in probing the strong interaction at unprecedented mass scales.

hep-ph

Recursive Reward Aggregation

In reinforcement learning (RL), aligning agent behavior with specific objectives typically requires careful design of the reward function, which can be challenging when the desired objectives are complex. In this work, we propose an alternative approach for flexible behavior alignment that eliminates the need to modify the reward function by selecting appropriate reward aggregation functions. By introducing an algebraic perspective on Markov decision processes (MDPs), we show that the Bellman equations naturally emerge from the recursive generation and aggregation of rewards, allowing for the generalization of the standard discounted sum to other recursive aggregations, such as discounted max and Sharpe ratio. Our approach applies to both deterministic and stochastic settings and integrates seamlessly with value-based and actor-critic algorithms. Experimental results demonstrate that our approach effectively optimizes diverse objectives, highlighting its versatility and potential for real-world applications.

cs.LG

Probing Yoctosecond Quantum Dynamics in Toponium Formation at Colliders

The formation of toponium, a bound state of top and anti-top quarks, provides an unprecedented system for investigating quantum state dynamics at ultrashort timescales. We explore two distinct phenomenological descriptions of this process: a 'wavelike' scenario emphasizing the role of quantum superposition at creation, and a 'particlelike' scenario where a finite formation time is governed by relativistic causality. Both descriptions are compatible with the principles of relativistic quantum field theory. We propose to distinguish these scenarios by exploiting the top quark's intrinsic lifetime ($τ_t \sim 5.02 \times 10^{-25}$\,s) as a quantum chronometer. We simulate the cross-section ratio $R_b = σ(e^+e^- \to b\bar{b})/\sum_q σ(e^+e^- \to q\bar{q})$ ($q = u,d,s,c,b$) near $\sqrt{s} = 343$\,GeV at future lepton colliders (CEPC/FCC-ee). These distinct scenarios yield observable $R_b$ profiles, enabling $>5σ$ discrimination with 1500\,fb$^{-1}$ of data. Preliminary LHC data provide independent $2$--$3σ$ support. This framework establishes a collider-based method to explore time-dependent quantum phenomena in particle production with yoctosecond ($10^{-24}$\,s) resolution.

hep-ph

$η_c$ leading-twist distribution amplitude and the $B_c \to η_c\ell\barν_\ell$ semileptonic decays using QCD Sum Rules

In this paper, we investigate the semileptonic decays $B_c \to η_c\ell\barν_\ell$ using the quantum chromodynamics(QCD) sum rules within the framework of Standard Model (SM). We further explore the potential to probe signatures of new Physics (NP) beyond the SM through these decays. First, we derive the $ξ$-moments $\langleξ_{2;η_c}^{n}\rangle$ of the $η_c$-meson leading-twist distribution amplitude $ϕ_{2;η_c}$ using the QCD sum rules within the background field theory. Considering contributions from the vacuum condensates up to dimension-six, the first two nonzero $ξ$-moments at the scale of $4$ GeV are found to be $\langleξ_{2;η_c}^{2}\rangle = 0.103^{+0.009}_{-0.009}$ and $\langleξ_{2;η_c}^{4}\rangle = 0.031^{+0.003}_{-0.003}$. Using these moments, we then fix the Gegenbauer expansion series of $ϕ_{2;η_c}$ and apply it to compute the $B_c \to η_c$ transition form factors (TFFs) using QCD light cone sum rules. Second, we extrapolate those TFFs to physically allowable $q^2$-range via a simplified series expansion, and we obtain $R_{η_c}|_{\rm SM} = 0.308^{+0.084}_{-0.062}$. Furthermore, we explore the potential impacts of various NP scenarios on $R_{η_c}$. Specifically, we compute the forward-backward asymmetry $\mathcal{A}_{\rm FB}({q^2})$, the convexity parameter $\mathcal{C}_F^τ({q^2})$, and the longitudinal and transverse polarizations $\mathcal{P}_L ({q^2})$ and $\mathcal{P}_T ({q^2})$ for $B_c \to η_c$ transitions within both the SM and two types of NP scenarios. Our results contribute to a deeper understanding of $B_c$-meson semileptonic decays and provide insights into the search for the NP beyond the SM.

hep-ph

Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update

We study the stochastic linear bandits with heavy-tailed noise. Two principled strategies for handling heavy-tailed noise, truncation and median-of-means, have been introduced to heavy-tailed bandits. Nonetheless, these methods rely on specific noise assumptions or bandit structures, limiting their applicability to general settings. The recent work [Huang et al.2024] develops a soft truncation method via the adaptive Huber regression to address these limitations. However, their method suffers undesired computational costs: it requires storing all historical data and performing a full pass over these data at each round. In this paper, we propose a \emph{one-pass} algorithm based on the online mirror descent framework. Our method updates using only current data at each round, reducing the per-round computational cost from $\mathcal{O}(t \log T)$ to $\mathcal{O}(1)$ with respect to current round $t$ and the time horizon $T$, and achieves a near-optimal and variance-aware regret of order $\widetilde{\mathcal{O}}\big(d T^{\frac{1-ε}{2(1+ε)}} \sqrt{\sum_{t=1}^T ν_t^2} + d T^{\frac{1-ε}{2(1+ε)}}\big)$ where $d$ is the dimension and $ν_t^{1+ε}$ is the $(1+ε)$-th central moment of reward at round $t$.

cs.LG

Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability

Non-stationary online learning has drawn much attention in recent years. Despite considerable progress, dynamic regret minimization has primarily focused on convex functions, leaving the functions with stronger curvature (e.g., squared or logistic loss) underexplored. In this work, we address this gap by showing that the regret can be substantially improved by leveraging the concept of mixability, a property that generalizes exp-concavity to effectively capture loss curvature. Let $d$ denote the dimensionality and $P_T$ the path length of comparators that reflects the environmental non-stationarity. We demonstrate that an exponential-weight method with fixed-share updates achieves an $\mathcal{O}(d T^{1/3} P_T^{2/3} \log T)$ dynamic regret for mixable losses, improving upon the best-known $\mathcal{O}(d^{10/3} T^{1/3} P_T^{2/3} \log T)$ result (Baby and Wang, 2021) in $d$. More importantly, this improvement arises from a simple yet powerful analytical framework that exploits the mixability, which avoids the Karush-Kuhn-Tucker-based analysis required by existing work.

cs.LG

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learning from human feedback (RLHF), and offline preference optimization, such as direct preference optimization. Conventional approaches typically assume accurate annotations. However, real-world preference data often contains noise due to human errors or biases. We propose a principled framework for robust policy optimization under noisy preferences, viewing reward modeling as a classification problem. This allows us to leverage symmetric losses, known for their robustness to label noise in classification, leading to our Symmetric Preference Optimization (SymPO) method. We prove that symmetric losses enable successful policy optimization even under noisy labels, as the resulting reward remains rank-preserving -- a property sufficient for policy improvement. Experiments on synthetic and real-world tasks demonstrate the effectiveness of SymPO.

cs.LG

Toponium: the smallest bound state and simplest hadron in quantum mechanics

We explore toponium, the smallest known quantum bound state of a top quark and its antiparticle, bound by the strong force. With a Bohr radius of $8\times 10^{-18}$~m and a lifetime of $2.5 \times 10^{-25}$~s, toponium uniquely probes microphysics. Unlike all other hadrons, it is governed by ultraviolet freedom. This distinction offers novel insights into quantum chromodynamics. Our analysis reveals a toponium signal exceeding $5σ$ in the distribution of the cross section ratio between $e^+e^- \rightarrow b\bar{b}$ and $e^+e^- \rightarrow q\bar{q}$ ($q=b,c,s,d,u$), based on 400~fb$^{-1}$ of data collected at $\sqrt{s}\approx 341~{\rm GeV}$. This discovery enables a top quark mass measurement with an uncertainty reduced by a factor of ten compared to current precision levels. Moreover, this method improves the systematic uncertainty by at least a factor of 2.7 compared to any other possible methods.

hep-ph

Toponium: Implementation of a toponium model in FeynRules

Toponium -- a bound state of the top-antitop pair ($t\bar{t}$) -- emerges as the smallest and simplest hadronic system in QCD, with an ultrashort lifetime ($τ_t \sim 2.5\times 10^{-25}$~s) and a femtometer-scale Bohr radius ($r_{\text{Bohr}} \sim 7\times 10^{-18}$~m). We present a computational framework extending the Standard Model (SM) with two S-wave toponium states: a spin-singlet $η_t$ ($J^{PC}=0^{-+}$) and a spin-triplet $J_t$ ($J^{PC}=1^{--}$). Using nonrelativistic QCD (NRQCD) and a Coulomb potential, we derived couplings to SM particles (gluons, electroweak bosons, Higgs boson, and fermion pairs) and implemented the Lagrangian in FeynRules, generating FeynArts, MadGraph, and WHIZARD models for collider simulations. Key results include dominant decay channels ($η_t \to gg/ZH$, $J_t \to W^+W^-/b\bar{b}$) and leading order (LO) cross sections for $pp \to η_t(nS) \to {\rm non-}t\bar{t}$ ({66 fb} at 13 TeV). The model avoids double-counting artifacts by excluding direct $t\bar{t}$ couplings, thereby ensuring consistency with perturbative QCD. This work establishes a complete pipeline for precision toponium studies, bridging NRQCD, collider phenomenology, and tests of SM validity at future lepton colliders (e.g., CEPC, FCC-ee, muon colliders) and the LHC. This provides the first publicly available UFO model for toponium, enabling direct integration with MadGraph and WHIZARD for simulations.

hep-ph

Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation

We study a new class of MDPs that employs multinomial logit (MNL) function approximation to ensure valid probability distributions over the state space. Despite its significant benefits, incorporating the non-linear function raises substantial challenges in both statistical and computational efficiency. The best-known result of Hwang and Oh [2023] has achieved an $\widetilde{\mathcal{O}}(κ^{-1}dH^2\sqrt{K})$ regret upper bound, where $κ$ is a problem-dependent quantity, $d$ is the feature dimension, $H$ is the episode length, and $K$ is the number of episodes. However, we observe that $κ^{-1}$ exhibits polynomial dependence on the number of reachable states, which can be as large as the state space size in the worst case and thus undermines the motivation for function approximation. Additionally, their method requires storing all historical data and the time complexity scales linearly with the episode count, which is computationally expensive. In this work, we propose a statistically efficient algorithm that achieves a regret of $\widetilde{\mathcal{O}}(dH^2\sqrt{K} + κ^{-1}d^2H^2)$, eliminating the dependence on $κ^{-1}$ in the dominant term for the first time. We then address the computational challenges by introducing an enhanced algorithm that achieves the same regret guarantee but with only constant cost. Finally, we establish the first lower bound for this problem, justifying the optimality of our results in $d$ and $K$.

cs.LG

Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More Practical

Complementary-label learning is a weakly supervised learning problem in which each training example is associated with one or multiple complementary labels indicating the classes to which it does not belong. Existing consistent approaches have relied on the uniform distribution assumption to model the generation of complementary labels, or on an ordinary-label training set to estimate the transition matrix in non-uniform cases. However, either condition may not be satisfied in real-world scenarios. In this paper, we propose a novel consistent approach that does not rely on these conditions. Inspired by the positive-unlabeled (PU) learning literature, we propose an unbiased risk estimator based on the Selected-Completely-at-Random assumption for complementary-label learning. We then introduce a risk-correction approach to address overfitting problems. Furthermore, we find that complementary-label learning can be expressed as a set of negative-unlabeled binary classification problems when using the one-versus-rest strategy. Extensive experimental results on both synthetic and real-world benchmark datasets validate the superiority of our proposed approach over state-of-the-art methods.

cs.LG

Exploratory Machine Learning with Unknown Unknowns

In conventional supervised learning, a training dataset is given with ground-truth labels from a known label set, and the learned model will classify unseen instances to known labels. This paper studies a new problem setting in which there are unknown classes in the training data misperceived as other labels, and thus their existence appears unknown from the given supervision. We attribute the unknown unknowns to the fact that the training dataset is badly advised by the incompletely perceived label space due to the insufficient feature information. To this end, we propose the exploratory machine learning, which examines and investigates training data by actively augmenting the feature space to discover potentially hidden classes. Our method consists of three ingredients including rejection model, feature exploration, and model cascade. We provide theoretical analysis to justify its superiority, and validate the effectiveness on both synthetic and real datasets.

cs.LG