SearcharxivSearch

arXiv subjects

Mao Zhang

Publications and source records attributed to Mao Zhang.

11 recordsLinked to original sources

DREAM Technical Report

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.

cs.IR

RecGPT-V3 Technical Report

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead. We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's "Guess What You Like" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.

cs.IR

RecGPT-V2 Technical Report

Large language models (LLMs) have demonstrated remarkable potential in transforming recommender systems from implicit behavioral pattern matching to explicit intent reasoning. While RecGPT-V1 successfully pioneered this paradigm by integrating LLM-based reasoning into user interest mining and item tag prediction, it suffers from four fundamental limitations: (1) computational inefficiency and cognitive redundancy across multiple reasoning routes; (2) insufficient explanation diversity in fixed-template generation; (3) limited generalization under supervised learning paradigms; and (4) simplistic outcome-focused evaluation that fails to match human standards. To address these challenges, we present RecGPT-V2 with four key innovations. First, a Hierarchical Multi-Agent System restructures intent reasoning through coordinated collaboration, eliminating cognitive duplication while enabling diverse intent coverage. Combined with Hybrid Representation Inference that compresses user-behavior contexts, our framework reduces GPU consumption by 60% and improves exclusive recall from 9.39% to 10.99%. Second, a Meta-Prompting framework dynamically generates contextually adaptive prompts, improving explanation diversity by +7.3%. Third, constrained reinforcement learning mitigates multi-reward conflicts, achieving +24.1% improvement in tag prediction and +13.0% in explanation acceptance. Fourth, an Agent-as-a-Judge framework decomposes assessment into multi-step reasoning, improving human preference alignment. Online A/B tests on Taobao demonstrate significant improvements: +2.98% CTR, +3.71% IPV, +2.19% TV, and +11.46% NER. RecGPT-V2 establishes both the technical feasibility and commercial viability of deploying LLM-powered intent reasoning at scale, bridging the gap between cognitive exploration and industrial utility.

cs.IR

RecGPT Technical Report

Recommender systems are among the most impactful applications of artificial intelligence, serving as critical infrastructure connecting users, merchants, and platforms. However, most current industrial systems remain heavily reliant on historical co-occurrence patterns and log-fitting objectives, i.e., optimizing for past user interactions without explicitly modeling user intent. This log-fitting approach often leads to overfitting to narrow historical preferences, failing to capture users' evolving and latent interests. As a result, it reinforces filter bubbles and long-tail phenomena, ultimately harming user experience and threatening the sustainability of the whole recommendation ecosystem. To address these challenges, we rethink the overall design paradigm of recommender systems and propose RecGPT, a next-generation framework that places user intent at the center of the recommendation pipeline. By integrating large language models (LLMs) into key stages of user interest mining, item retrieval, and explanation generation, RecGPT transforms log-fitting recommendation into an intent-centric process. To effectively align general-purpose LLMs to the above domain-specific recommendation tasks at scale, RecGPT incorporates a multi-stage training paradigm, which integrates reasoning-enhanced pre-alignment and self-training evolution, guided by a Human-LLM cooperative judge system. Currently, RecGPT has been fully deployed on the Taobao App. Online experiments demonstrate that RecGPT achieves consistent performance gains across stakeholders: users benefit from increased content diversity and satisfaction, merchants and the platform gain greater exposure and conversions. These comprehensive improvement results across all stakeholders validates that LLM-driven, intent-centric design can foster a more sustainable and mutually beneficial recommendation ecosystem.

cs.IR

Bursting Filter Bubble: Enhancing Serendipity Recommendations with Aligned Large Language Models

Recommender systems (RSs) often suffer from the feedback loop phenomenon, e.g., RSs are trained on data biased by their recommendations. This leads to the filter bubble effect that reinforces homogeneous content and reduces user satisfaction. To this end, serendipity recommendations, which offer unexpected yet relevant items, are proposed. Recently, large language models (LLMs) have shown potential in serendipity prediction due to their extensive world knowledge and reasoning capabilities. However, they still face challenges in aligning serendipity judgments with human assessments, handling long user behavior sequences, and meeting the latency requirements of industrial RSs. To address these issues, we propose SERAL (Serendipity Recommendations with Aligned Large Language Models), a framework comprising three stages: (1) Cognition Profile Generation to compress user behavior into multi-level profiles; (2) SerenGPT Alignment to align serendipity judgments with human preferences using enriched training data; and (3) Nearline Adaptation to integrate SerenGPT into industrial RSs pipelines efficiently. Online experiments demonstrate that SERAL improves exposure ratio (PVR), clicks, and transactions of serendipitous items by 5.7%, 29.56%, and 27.6%, enhancing user experience without much impact on overall revenue. Now, it has been fully deployed in the "Guess What You Like" of the Taobao App homepage.

cs.IR

Quantum speed limit for complex dynamics

Quantum speed limit focuses on the minimum time scale for a fixed mission and hence is important in quantum information where fast dynamics is usually beneficial. Most existing tools for the depiction of quantum speed limit are the lower-bound-type tools, which are in fact difficult to reveal the true minimum time, especially for many-body systems or complex dynamics. Therefore, the evaluation of this true minimum time in these scenarios is still an unsolved problem. Hereby we propose a three-step (classification-regression-calibration) methodology based on machine learning to evaluate the true minimum time in complex dynamics. Moreover, the analytical expression of the true minimum time is also provided for the time-dependent Hamiltonians with time-independent eigenstates.

quant-ph

QuanEstimation: An open-source toolkit for quantum parameter estimation

Quantum parameter estimation promises a high-precision measurement in theory, however, how to design the optimal scheme in a specific scenario, especially under a practical condition, is still a serious problem that needs to be solved case by case due to the existence of multiple mathematical bounds and optimization methods. Depending on the scenario considered, different bounds may be more or less suitable, both in terms of computational complexity and the tightness of the bound itself. At the same time, the metrological schemes provided by different optimization methods need to be tested against realization complexity, robustness, etc. Hence, a comprehensive toolkit containing various bounds and optimization methods is essential for the scheme design in quantum metrology. To fill this vacancy, here we present a Python-Julia-based open-source toolkit for quantum parameter estimation, which includes many well-used mathematical bounds and optimization methods. Utilizing this toolkit, all procedures in the scheme design, such as the optimizations of the probe state, control and measurement, can be readily and efficiently performed.

quant-ph

Optimal Scheme for Quantum Metrology

Quantum metrology can achieve far better precision than classical metrology, and is one of the most important applications of quantum technologies in the real world. To attain the highest precision promised by quantum metrology, all steps of the schemes need to be optimized, which include the state preparation, parametrization, and measurement. Here the recent progresses on the optimization of these steps, which are essential for the identification and achievement of the ultimate precision limit in quantum metrology, are reviewed. It is hoped this provides a useful reference for the researchers in quantum metrology and related fields.

quant-ph

Quantum metrology with precision reaching beyond-$1/N$ scaling through $N$-probe entanglement generating interactions

Nonlinear interactions are recognized as potential resources for quantum metrology, facilitating parameter estimation precisions that scale as the exponential Heisenberg limit of $2^{-N}$. We explore such nonlinearity and propose an associated quantum measurement scenario based on the nonlinear interaction of $N$-probe entanglement generating form. This scenario provides an enhanced precision scaling of $D^{-N}/(N-1)!$ with $D > 2$ a tunable parameter. In addition, it can be readily implemented in a variety of experimental platforms and applied to measurements of a wide range of quantities, including local gravitational acceleration $g$, magnetic field, and its higher-order gradients.

quant-ph

Generation and storage of spin squeezing via learning-assisted optimal control

The generation and storage of spin squeezing is an attracting topic in quantum metrology and the foundations of quantum mechanics. The major models to realize the spin squeezing are the one- and two-axis twisting models. Here, we consider a collective spin system coupled to a bosonic field, and show that proper constant-value controls in this model can simulate the dynamical behaviors of these two models. More interestingly, a better performance of squeezing can be obtained when the control is time-varying, which is generated via a reinforcement learning algorithm. However, this advantage becomes limited if the collective noise is involved. To deal with it, we propose a four-step strategy for the construction of a new type of combined controls, which include both constant-value and time-varying controls, but performed at different time intervals. Compared to the full time-varying controls, the combined controls not only give a comparable minimum value of the squeezing parameter over time, but also provides a better lifetime and larger full amount of squeezing. Moreover, the amplitude form of a combined control is simpler and more stable than the full time-varying control. Therefore, our scheme is very promising to be applied in practice to improve the generation and storage performance of squeezing.

quant-ph

Operational definition of a quantum speed limit

The quantum speed limit is a fundamental concept in quantum mechanics, which aims at finding the minimum time scale or the maximum dynamical speed for some fixed targets. In a large number of studies in this field, the construction of valid bounds for the evolution time is always the core mission, yet the physics behind it and some fundamental questions like which states can really fulfill the target, are ignored. Understanding the physics behind the bounds is at least as important as constructing attainable bounds. Here we provide an operational approach for the definition of the quantum speed limit, which utilizes the set of states that can fulfill the target to define the speed limit. Its performances in various scenarios have been investigated. For time-independent Hamiltonians, it is inverse-proportional to the difference between the highest and lowest energies. The fact that its attainability does not require a zero ground-state energy suggests it can be used as an indicator of quantum phase transitions. For time-dependent Hamiltonians, it is shown that contrary to the results given by existing bounds, the true speed limit should be independent of the time. Moreover, in the case of spontaneous emission, we find a counterintuitive phenomenon that a lousy purity can benefit the reduction of the quantum speed limit.

quant-ph