Searcharxiv⌕ Search

arXiv subjects

Zhicheng Zhu

Publications and source records attributed to Zhicheng Zhu.

13 recordsLinked to original sources

GR2 Technical Report

Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking -- where the final re-ranking step disproportionately shapes user engagement and downstream performance, particularly for carousel and grid display formats. Despite growing enthusiasm for Large Language Models (LLMs) in recommendation, three gaps hinder industrial adoption: (1) most efforts target retrieval and ranking, leaving re-ranking -- the stage closest to the final user experience -- largely underexplored; (2) LLMs are typically deployed zero-shot or via supervised fine-tuning, underutilizing the reasoning capabilities unlocked by reinforcement learning (RL) on verifiable rewards; (3) deployed catalogs index billions of items with non-semantic identifiers that lie outside any base-LLM vocabulary. We present GR2 (Generative Reasoning Re-Ranker), an end-to-end framework that combines (i) mid-training on semantic IDs produced by a tokenizer with >=99% uniqueness, (ii) reasoning-trace distilled from a stronger teacher via targeted prompting and rejection sampling, and (iii) RL with verifiable rewards purpose-built for re-ranking. To make GR2 resource-viable, we further (iv) introduce a context compressor that amortizes training cost, On-Policy Distillation (OPD) as a scalable alternative to SFT -- which we find collapses at industrial scale -- and reasoning distillation for low-latency serving. GR2 delivers +18.7% R@1, +7.1% R@3, and +9.6% N@3 over legacy baselines on industrial-scale traffic. We further find that reward design is critical in re-ranking: LLMs often hack rewards by preserving the incoming order or exploiting position bias, motivating conditional verifiable rewards as essential industrial components.

cs.IR↗

Fertility fibres and coproduct coefficients in the LOT Hopf algebra

We study fibres of the fertility map $Φ$ from decorated rooted trees to decorated multi-index monomials. For a multi-index $\mathbf{k}$ of weight $-1$, the fibre $\mathcal F_{\mathbf{k}}=\{\,t:Φ(t)=\xx^{\mathbf{k}}\,\}$ consists of all rooted trees with decoration--fertility profile $\mathbf{k}$. We consider its ordinary cardinality $F_{\mathbf{k}}$, its symmetry-weighted cardinality $W_{\mathbf{k}}$, and the coefficient mass $J_{\mathbf{k}}$ appearing in the tree expansion of the transposed embedding $\jmath$. We obtain an explicit formula and a functional equation for the weighted counts, and an exact multiset recursion together with a cycle-index functional equation for the ordinary counts. We also introduce coefficient generating functions for the lowering derivation $\bar\partial$, derive recursive and transport-array formulas for the corresponding coefficients, and use them to refine the admissible-cut formula for the coproduct in the LOT Hopf algebra.

math.CO↗

Aromatic and clumped multi-indices: algebraic structure and Hopf embeddings

Butcher forests extend naturally into aromatic and clumped forests and play a fundamental role in the numerical analysis of volume-preserving methods. The design of general volume-preserving methods is a challenging open problem, and recent attempts showed progress on specific dynamics. We introduce aromatic and clumped multi-indices, that are algebraic objects that simplify the study of volume-preservation to the one-dimensional setting, while retaining much of the structure (in stark opposition to standard multi-indices). We provide their algebraic structure of pre-Lie-Rinehart algebra, Hopf algebroid, and Hopf algebra, apply them in numerical analysis, and generalise to the aromatic context the Hopf embedding from multi-indices to the BCK Hopf algebra.

math.CO↗

Banach fixed point and flow approach for rough analysis

In this paper, we show that the main algebraic assumption required to perform a fixed point argument for rough differential equations implies the algebraic assumption for the Bailleul flow approach. This assumption requires that the rough path associated with the equation is given by a Hopf algebra whose coproduct admits a cocycle and has a tree-like basis. We show that the Hopf algebra of multi-indices does not satisfy the cocycle condition. This is a rigorous result on the impossibility, observed in practice, of performing a fixed point argument for multi-indices rough paths and multi-indices in Regularity Structures.

math.PR↗

BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training

Large Language Models (LLMs) have become a cornerstone of modern AI, driving breakthroughs in natural language processing and expanding into multimodal jobs involving images, audio, and video. As with most computational software, it is important to distinguish between ordinary runtime performance and startup overhead. Prior research has focused on runtime performance: improving training efficiency and stability. This work focuses instead on the increasingly critical issue of startup overhead in training: the delay before training jobs begin execution. Startup overhead is particularly important in large, industrial-scale LLMs, where failures occur more frequently and multiple teams operate in iterative update-debug cycles. In one of our training clusters, more than 3.5% of GPU time is wasted due to startup overhead alone. In this work, we present the first in-depth characterization of LLM training startup overhead based on real production data. We analyze the components of startup cost, quantify its direct impact, and examine how it scales with job size. These insights motivate the design of Bootseer, a system-level optimization framework that addresses three primary startup bottlenecks: (a) container image loading, (b) runtime dependency installation, and (c) model checkpoint resumption. To mitigate these bottlenecks, Bootseer introduces three techniques: (a) hot block record-and-prefetch, (b) dependency snapshotting, and (c) striped HDFS-FUSE. Bootseer has been deployed in a production environment and evaluated on real LLM training workloads, demonstrating a 50% reduction in startup overhead.

cs.LG↗

Scaling Reinforcement Learning for Content Moderation with Large Language Models

Content moderation at scale remains one of the most pressing challenges in today's digital ecosystem, where billions of user- and AI-generated artifacts must be continuously evaluated for policy violations. Although recent advances in large language models (LLMs) have demonstrated strong potential for policy-grounded moderation, the practical challenges of training these systems to achieve expert-level accuracy in real-world settings remain largely unexplored, particularly in regimes characterized by label sparsity, evolving policy definitions, and the need for nuanced reasoning beyond shallow pattern matching. In this work, we present a comprehensive empirical investigation of scaling reinforcement learning (RL) for content classification, systematically evaluating multiple RL training recipes and reward-shaping strategies-including verifiable rewards and LLM-as-judge frameworks-to transform general-purpose language models into specialized, policy-aligned classifiers across three real-world content moderation tasks. Our findings provide actionable insights for industrial-scale moderation systems, demonstrating that RL exhibits sigmoid-like scaling behavior in which performance improves smoothly with increased training data, rollouts, and optimization steps before gradually saturating. Moreover, we show that RL substantially improves performance on tasks requiring complex policy-grounded reasoning while achieving up to 100x higher data efficiency than supervised fine-tuning, making it particularly effective in domains where expert annotations are scarce or costly.

cs.AI↗

Controlled rough paths: a general Hopf-algebraic setting

We set up controlled rough paths for a class of combinatorial Hopf algebras, encompassing shuffle, Butcher-Connes-Kreimer and Munthe-Kaas--Wright Hopf algebras. The class of controls we consider encompasses both Hölder continuous paths and (not necessarily continuous) paths with bounded $p$-variation. We prove existence and uniqueness of the solution of a lifted initial value problem in this general setting by applying the fixed point method in a suitable Banach space of controlled rough paths, and we prove a universal limit theorem addressing the robustness of the solution with respect to the parameters and the initial condition.

math.PR↗

Free Novikov algebras and the Hopf algebra of decorated multi-indices

We propose a combinatorial formula for the coproduct in a Hopf algebra of decorated multi-indices that recently appeared in the literature, which can be briefly described as the graded dual of the enveloping algebra of the free Novikov algebra generated by the set of decorations. Similarly to what happens for the Hopf algebra of rooted forests, the formula can be written in terms of admissible cuts. We also prove a combinatorial formula for the extraction-contraction coproduct for undecorated multi-indices, in terms of a suitable notion of covering subforest.

math.CO↗

New operated polynomial identities and Gröbner-Shirshov bases

Quite recently, Bremner et al. introduced a new approach to Rota's Classification Problem and classified some (new) operated polynomial identities. In this paper, we prove that all operated polynomial identities classified by Bremner et al. are Gröbner-Shirshov.

math.RA↗

Robust Remanufacturing Planning with Parameter Uncertainty

We consider the problem of remanufacturing planning in the presence of statistical estimation errors. Determining the optimal remanufacturing timing, first and foremost, requires modeling of the state transitions of a system. The estimation of these probabilities, however, often suffers from data inadequacy and is far from accurate, resulting in serious degradation in performance. To mitigate the impacts of the uncertainty in transition probabilities, we develop a novel data-driven modeling framework for remanufacturing planning in which decision makers can remain robust with respect to statistical estimation errors. We model the remanufacturing planning problem as a robust Markov decision process, and construct ambiguity sets that contain the true transition probability distributions with high confidence. We further establish structural properties of optimal robust policies and insights for remanufacturing planning. A computational study on the NASA turbofan engine shows that our data-driven decision framework consistently yields better worst-case performances and higher reliability of the performance guarantee

math.OC↗

Condition-based Maintenance for Multi-component Systems:Modeling, Structural Properties, and Algorithms

Condition-based maintenance (CBM) is an effective maintenance strategy to improve system performance while lowering operating and maintenance costs. Real-world systems typically consist of a large number of components with various interactions between components. However, existing studies on CBM focus on single-component systems. Multi-component condition-based maintenance, which joins the components' stochastic degradation processes and the combinatorial maintenance grouping problem, remains an open issue in the literature. In this paper, we study the CBM optimization problem for multi-component systems. We first develop a multi-stage stochastic integer model with the objective of minimizing the total maintenance cost over a finite planning horizon. We then investigate the structural properties of a two-stage model. Based on the structural properties, two efficient algorithms are designed to solve the two-stage model. Algorithm 1 solves the problem to its optimality and Algorithm 2 heuristically searches for high-quality solutions based on Algorithm 1. Our computational studies show that Algorithm 1 obtains optimal solutions in a reasonable amount of time and Algorithm 2 can find high-quality solutions quickly. The multi-stage problem is solved using a rolling horizon approach based on the algorithms for the two-stage problem.

math.OC↗

Multi-component Maintenance Optimization: A Stochastic Programming Approach

Maintenance optimization has been extensively studied in the past decades. However, most of the existing maintenance models focus on single-component systems and are not applicable for complex systems consisting of multiple components, due to various interactions between the components. Multi-component maintenance optimization problem, which joins the stochastic processes regarding the failures of the components with the combinatorial problems regarding the grouping of maintenance activities, is challenging in both modeling and solution techniques, and has remained as an open issue in the literature. In this paper, we study the multi-component maintenance problem over a finite planning horizon and formulate the problem as a multi-stage stochastic integer program with decision-dependent uncertainty. There is a lack of general efficient methods to solve this type of problem. To address this challenge, we use an alternative approach to model the underlying failure process and develop a novel two-stage model without decision-dependent uncertainty. Structural properties of the two-stage problem are investigated, and a progressive-hedging-based heuristic is developed based on the structural properties. Our heuristic algorithm demonstrates a significantly improved capacity in handling practically large-size two-stage problems comparing to three conventional methods for stochastic integer programming, and solving the two-stage model by our heuristic in a rolling horizon provides a good approximation of the multi-stage problem. The heuristic is further benchmarked with a dynamic programming approach commonly adopted in the literature. Numerical results show that our heuristic can lead to significant cost savings compared with the benchmark approach.

math.OC↗