SearcharxivSearch

arXiv subjects

Huihui Cheng

Publications and source records attributed to Huihui Cheng.

6 recordsLinked to original sources

SlimPer: Make Personalization Model Slim and Smart

Transformer-style architectures are increasingly adopted for industrial recommendation systems, yet they inherit a design premise misaligned with the task: generative models rely on per-token autoregressive prediction, which justifies maintaining large intermediate tensors that scale with sequence length. In contrast, recommendation systems produce a single set of relevance scores for each pair without token-level supervision. Leveraging this observation, we propose SlimPer, which reformulates personalized ranking as iterative refinement of a compact, unified knowledge base. At each layer, the model selectively queries raw multi-modal user-side tokens, computes explicit relevance matching scores, and refines the knowledge base, all in O(N) per-layer cost with a fixed-size intermediate representation. As a result, model depth is decoupled from user history length, enabling deeper relevance understanding without proportional growth in compute or memory; request-only optimization further trims memory by sharing a single copy of user-side tokens across all candidate items. SlimPer unifies sparse, dense, and sequence features within a single backbone and provides inherent interpretability through its attention mechanism. Deployed on Instagram Reels and Feed, SlimPer yields measurable improvements in user engagement while streamlining the overall system and enabling effective modeling of 10k+ fine-grained user history events.

cs.IR

Target-Aware Early Stage Ranking

Early Stage Ranking (ESR) in large-scale recommendation systems is dominated by ''user--item decoupling'' Two Tower architectures, which scale efficiently but cannot capture fine-grained, target-aware user--item interactions directly. We propose Target-Aware Early Stage Ranking (TESR), which augments the Two Tower with a Mixture of Attention (MoA) module trained as a request-level sequence modeling over user history. MoA combines (i) Hard Matching Attention (HMA) to capture explicit categorical-ID level overlap signals between user history and candidate item, (ii) target-aware HSTU attention for implicit affinities conditioned on the candidate, and (iii) target dependent and independent cross-attention for symmetric user-item contextualization. On top of this, a Multi-Logit Parameterized Gating (MLPG) head amplifies these signals at scoring time. To keep latency within ESR budgets, we co-design the architecture with FP8 quantization, custom kernels, and a Torch Inductor compilation path. On a production deployment, TESR delivers consistent offline NE wins and online topline gains, and is, to our knowledge, the first deployment of full target-aware attention sequence modeling in an ESR stage at this scale.

cs.LG

Request-Only Optimization for Recommendation Systems

Deep Learning Recommendation Models (DLRMs) represent one of the largest machine learning applications on the planet. Industry-scale DLRMs are trained with petabytes of recommendation data to serve billions of users every day. To utilize the rich user signals in the long user history, DLRMs have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems. In this paper, we present a Request-Only Optimizations (ROO) training and modeling paradigm. ROO simultaneously improves the storage and training efficiency as well as the model quality of recommendation systems. We holistically approach this challenge through co-designing data (i.e., request-only data), infrastructure (i.e., request-only based data processing pipeline), and model architecture (i.e., request-only neural architectures). Our ROO training and modeling paradigm treats a user request as a unit of the training data. Compared with the established practice of treating a user impression as a unit, our new design achieves native feature deduplication in data logging, consequently saving data storage. Second, by de-duplicating computations and communications across multiple impressions in a request, this new paradigm enables highly scaled-up neural network architectures to better capture user interest signals, such as Generative Recommenders (GRs) and other request-only friendly architectures.

cs.IR

Asymptotic Normality of the Largest Eigenvalue for Noncentral Sample Covariance Matrices

Let $X$ be a $p\times n$ independent identically distributed real Gaussian matrix with positive mean $\mu $ and variance $\sigma^2$ entries. The goal of this paper is to investigate the largest eigenvalue of the noncentral sample covariance matrix $W=XX^{T}/n$, when the dimension $p$ and the sample size $n$ both grow to infinity with the limit $p/n=c\,(0<c<\infty)$. Utilizing the von Mises iteration method, we derive an approximation of the largest eigenvalue $\lambda_{1}(W)$ and show that $\lambda_{1}(W)$ asymptotically has a normal distribution with expectation $p\mu^2+(1+c)\sigma^2$ and variance $4c\mu^2\sigma^2$.

math.PR

1.06 μm Q-switched ytterbium-doped fiber laser using few-layer topological insulator Bi2Se3 as a saturable absorber

Passive Q-switching of an ytterbium-doped fiber (YDF) laser with few-layer topological insulator (TI) is, to the best of our knowledge, experimentally demonstrated for the first time. The few-layer TI: Bi2Se3 (2-4 layer thickness) is fabricated by the liquid-phase exfoliation method, and has a low saturable optical intensity of 53 MW/cm2 measured by the Z-scan technique. The optical deposition technique is used to induce the few-layer TI in the solution onto a fiber ferrule for successfully constructing the fiber-integrated TI-based saturable absorber (SA). By inserting this SA into the YDF laser cavity, stable Q-switching operation at 1.06 μm is achieved. The Q-switched pulses have the shortest pulse duration of 1.95 μs, the maximum pulse energy of 17.9 nJ and a tunable pulse-repetition-rate from 8.3 to 29.1 kHz. Our results indicate that the TI as a SA is also available at 1 μm waveband, revealing its potential as another wavelength-independent SA (like graphene).

physics.optics

Exclusive $B \to V γ$ decays in the T2HDM

By employing the QCD factorization approach for the exclusive $B \to V γ$ decays, we calculated the new physics contributions to the branching ratios, CP asymmetries, isospin and U-spin symmetry breaking of $B \to K^*γ$ and $B \to ργ$ decays, induced by the charged Higgs penguin diagrams appeared in the top-quark two-Higgs-doublet model(T2HDM). Within the considered parameter space, we found that (a) a charged-Higgs boson with a mass larger than 300 GeV are always allowed by the date of $B \to V γ$ decay, and such lower limit on $\mhp$ are comparable with those obtained from the inclusive $B \to X_s γ$ decay; (b) the CP asymmetry of $B \to ργ$ in the T2HDM can be as large as 10% in magnitude and has a strong dependence on the angle $θ$ and the CKM angle $γ$; (c) the isospin symmetry breakings of $B \to V γ$ decays in the T2HDM are generally small in size: around 6% for $B \to K^* γ$ decay and less than 20% for $B \to ργ$ decay; and (d) the U-spin symmetry breaking $ΔU(K^*,ρ)$ in the T2HDM is also small in size, only about 8% of the branching ratio $\calb (B \to ρ^0 γ)$.

hep-ph