SearcharxivSearch

arXiv subjects

Yujie Luo

Publications and source records attributed to Yujie Luo.

At least 19 recordsLinked to original sources

DREAM Technical Report

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.

cs.IR

Log Calabi--Yau structure for endomorphisms on $\mathbf{P}^n$

Let $f:\mathbf{P}^n\to\mathbf{P}^n$ be a $q$-polarized endomorphism, where $q>1$, and let $R_f$ be its ramification divisor. We study the singularities of the ramification pair $(\mathbf{P}^n,R_f)$. We show that, for a general $f$, the pair $(\mathbf{P}^n,R_f)$ is log canonical. When $n=2$, we prove that there exists an integer $s\geq1$ such that the log canonical threshold $\mathrm{lct}(\mathbf{P}^2;R_{f^s})\geq1/(q^s-1)$. The passage to an iterate is necessary in general, and the lower bound is optimal. In particular, $(\mathbf{P}^2,R_{f^s}/(q^s-1))$ is a log Calabi--Yau pair, completing the proof of Gongyo's conjecture for smooth projective surfaces.

math.AG

Sharp Bounds for Totally Invariant Cycles of Projective Varieties

Let $X$ be a smooth projective variety, $f:X\to X$ an int-amplified endomorphism, and $L$ any ample line bundle on $X$. We prove that, in every codimension, the total degree of prime cycles that become totally invariant under an iterate of $f$ satisfies an explicit Hilbert-function bound on their total $L$-degree. In particular, when $X=\mathbf{P}^n$, the number of totally invariant prime $(n-r)$-cycle is bounded by $\binom{n+1}{r}$, and this bound is optimal.

math.AG

RecGPT-V3 Technical Report

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead. We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's "Guess What You Like" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.

cs.IR

Essential dimensions of polarized endomorphisms of certain algebraic surfaces

Let $f: X \to X$ be a polarized endomorphism of a smooth projective surface which is birationally ruled. We answer a question of Koll\'ar and Zhuang, in the affirmative, on the incompressibility of $f$, under the assumption that $f$ is Galois and an explicit lower bound of deg$(f)$ depending only on $X$. We also give examples showing the optimality of such a lower bound.

math.AG

ClinReadNet: A clinical reading-inspired network for low-dose abdominal CT image quality assessment

In abdominal CT imaging, developing a low-dose, no-reference image quality assessment (No-reference IQA) model that mimics doctors' reading habits for evaluating CT image quality has significant practical value. This paper proposes a novel deep learning-based framework, ClinReadNet, whose design aligns with the clinical reading logic of radiologists: first, it introduces the Sobel ordinal quality network (SOQN) module, which can simultaneously focus on edge details highly relevant to image quality and the quality distribution pattern of the entire image, accurately matching the clinical image-reading judgment habit of "considering both local details and overall context"; second, the framework integrates the (shifted) window multi-scale temperature multi-head self-attention ((S)W-MTMSA) module, which further replicates the radiologists' image-reading process of shifting from overall scanning to local focusing, and accurately locks in regions of interest through multi-sharpness attention; third, it designs the hierarchical ranked probability score (HRPS) loss function, which combines the dual logics of coarse classification and fine classification, while paying attention to the distance information between grading labels, effectively improving the performance of image quality assessment. Experiments conducted on the LDCTIQAG2023 dataset show that the proposed method achieves the current state-of-the-art (SOTA) performance: the values of Pearson's linear correlation coefficient (PLCC), Spearman's rank-order correlation coefficient (SROCC), and Kendall's rank-order correlation coefficient (KROCC) reach 0.9507, 0.9554, and 0.8629 respectively, with the sum of their absolute values (Score) being 2.7690, outperforming existing methods.

cs.CV

Galois self-covers of projective spaces and essential dimensions

We give the structure theorem of Galois self-covers $f: \mathbf{P}^n \to \mathbf{P}^n$. As an application, we show that the essential dimension of every such nontrivial cover attains its maximum possible value $n$. As another application, we prove that the pair $(\mathbf{P}^n, R_f/(q-1))$ is log Calabi-Yau as conjectured by Gongyo, where $R_f$ is the ramification divisor and we write $f^*\mathcal{O}(1) = \mathcal{O}(q)$.

math.AG

Exploring Autonomous Agentic Data Engineering for Model Specialization

Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data. Existing LLM-based data curation methods primarily rely on human-designed workflows, leaving it unexamined whether LLMs can autonomously execute an end-to-end data engineering pipeline for model specialization. We formalize Autonomous Agentic Data Engineering, a novel task designed to evaluate LLMs as autonomous data engineers that drive model specialization through end-to-end data curation. We frame data as an optimizable component and study agents that plan, generate, and iteratively optimize training data across multiple domains, guided by post-training performance improvement. Experiments show that autonomous LLM data engineers yield substantial gains, as GPT-5.2 constructs a training curriculum that improves a student model by 57.29%, entirely through iterative, agent-driven data adaptation. By illuminating both potential and bottlenecks, our study establishes autonomous data engineering as a measurable capability and charts a path toward agent-driven model specialization (Code will be released at https://github.com/zjunlp/DataAgent).

cs.CL

Adaptive and ultrabroadband thermal control with solid-state nanophotonic emitters

Managing the emission and absorption of thermal radiation is crucial for a wide range of technologies, from radiative cooling of buildings and vehicles to thermal regulation of satellites and future lunar and Mars habitats. Despite this universal and critical need, thermal emitters capable of adaptively modulating emissivity in a broadband, high-contrast, and fully solid-state manner remain elusive. Here, we leverage neural-network-guided photonic design to enable adaptive, solid-state thermal emitters based on chalcogenide phase-change materials capable of emissivity switching with extreme spectral contrast and bandwidth. These engineered nanophotonic emitters operate over a broad spectrum$-$from solar through thermal infrared$-$providing very low solar absorptivity while enabling switchable thermal infrared emissivity with high contrast. We experimentally demonstrate the core functionality of our approach in the space-like radiative environment in the stratosphere, observing a 31.5 {\deg}C temperature differential between the two solid-state phases of a simplified chalcogenide GeSbTe-225 thermal emitter. Our results point to even more significant capabilities, such as the potential to modulate >600 W/m$^2$ of radiative heat (at 100 {\deg}C) with minimal solar heating in the vacuum of space. The proposed nanophotonic solid-state adaptive emitter could provide high-power and high-speed heat modulation while requiring no power to maintain state, offering transformative capabilities for thermal control in dynamic radiative environments on Earth and in space.

physics.app-ph

LightThinker++: From Reasoning Compression to Memory Management

Large language models (LLMs) excel at complex reasoning, yet their efficiency is limited by the surging cognitive overhead of long thought traces. In this paper, we propose LightThinker, a method that enables LLMs to dynamically compress intermediate thoughts into compact semantic representations. However, static compression often struggles with complex reasoning where the irreversible loss of intermediate details can lead to logical bottlenecks. To address this, we evolve the framework into LightThinker++, introducing Explicit Adaptive Memory Management. This paradigm shifts to behavioral-level management by incorporating explicit memory primitives, supported by a specialized trajectory synthesis pipeline to train purposeful memory scheduling. Extensive experiments demonstrate the framework's versatility across three dimensions. (1) LightThinker reduces peak token usage by 70% and inference time by 26% with minimal accuracy loss. (2) In standard reasoning, LightThinker++ slashes peak token usage by 69.9% while yielding a +2.42% accuracy gain under the same context budget for maximum performance. (3) Most notably, in long-horizon agentic tasks, it maintains a stable footprint beyond 80 rounds (a 60%-70% reduction), achieving an average performance gain of 14.8% across different complex scenarios. Overall, our work provides a scalable direction for sustaining deep LLM reasoning over extended horizons with minimal overhead.

cs.CL

Field-Programmable Mobile Magneto-Photonic Metaparticles for Active Light Manipulation and Steering

Controlling the flow of light within complex and dynamic environments is essential for a wide range of applications, from deep-tissue imaging and optogenetics to precision phototherapy. Typically, such light flows are controlled using external optical systems requiring line-of-sight access or by embedded nanoparticle scatterers with limited directional control, underscoring the need for mobile photonic agents capable of actively delivering and steering light within complex media. Here, we present magneto-photonic metaparticles: mobile, magnetically actuated microstructures that integrate a magnetic core with a nanoimprinted photonic surface. This hybrid design merges the reconfigurability of photonic metasurfaces with the mobility of magnetic actuation, enabling programmable translation, rotation, and real-time beam steering in aqueous media. In a concept-proof demonstration, we realize polymeric metaparticles with embedded magnetic core and nanoimprinted surface that exhibit controlled locomotion and active, magnetically programmable beam steering. Our design approach further points to more sophisticated metaparticle designs with high-efficiency, polarization-insensitive light steering, compatible with a scalable, single-step nanoimprint process. The mobile magneto-photonic metaparticle platform combines metasurface-level optical control with magnetic mobility, offering a versatile and scalable approach for active photonic control in complex environments.

physics.optics

Can We Predict Before Executing Machine Learning Agents?

Autonomous machine learning agents have revolutionized scientific discovery, yet they remain constrained by a Generate-Execute-Feedback paradigm. Previous approaches suffer from a severe Execution Bottleneck, as hypothesis evaluation relies strictly on expensive physical execution. To bypass these physical constraints, we internalize execution priors to substitute costly runtime checks with instantaneous predictive reasoning, drawing inspiration from World Models. In this work, we formalize the task of Data-centric Solution Preference and construct a comprehensive corpus of 18,438 pairwise comparisons. We demonstrate that LLMs exhibit significant predictive capabilities when primed with a Verified Data Analysis Report, achieving 61.5% accuracy and robust confidence calibration. Finally, we instantiate this framework in FOREAGENT, an agent that employs a Predict-then-Verify loop, achieving a 6x acceleration in convergence while surpassing execution-based baselines by +6%. Our code and dataset are publicly available at https://github.com/zjunlp/predict-before-execute.

cs.CL

RecGPT-V2 Technical Report

Large language models (LLMs) have demonstrated remarkable potential in transforming recommender systems from implicit behavioral pattern matching to explicit intent reasoning. While RecGPT-V1 successfully pioneered this paradigm by integrating LLM-based reasoning into user interest mining and item tag prediction, it suffers from four fundamental limitations: (1) computational inefficiency and cognitive redundancy across multiple reasoning routes; (2) insufficient explanation diversity in fixed-template generation; (3) limited generalization under supervised learning paradigms; and (4) simplistic outcome-focused evaluation that fails to match human standards. To address these challenges, we present RecGPT-V2 with four key innovations. First, a Hierarchical Multi-Agent System restructures intent reasoning through coordinated collaboration, eliminating cognitive duplication while enabling diverse intent coverage. Combined with Hybrid Representation Inference that compresses user-behavior contexts, our framework reduces GPU consumption by 60% and improves exclusive recall from 9.39% to 10.99%. Second, a Meta-Prompting framework dynamically generates contextually adaptive prompts, improving explanation diversity by +7.3%. Third, constrained reinforcement learning mitigates multi-reward conflicts, achieving +24.1% improvement in tag prediction and +13.0% in explanation acceptance. Fourth, an Agent-as-a-Judge framework decomposes assessment into multi-step reasoning, improving human preference alignment. Online A/B tests on Taobao demonstrate significant improvements: +2.98% CTR, +3.71% IPV, +2.19% TV, and +11.46% NER. RecGPT-V2 establishes both the technical feasibility and commercial viability of deploying LLM-powered intent reasoning at scale, bridging the gap between cognitive exploration and industrial utility.

cs.IR

Essential dimensions of polarized endomorphisms of abelian varieties

Let $f$ be a polarized endomorphism of an abelian variety $A$. Koll\'ar and Zhuang asked whether the essential dimension $\mathrm{ed}(f)$ equals $\mathrm{dim}(A)$. We provide counterexamples to this question. Instead, we prove that, under the hypothesis that every subtorus of $A$ is $f$-preperiodic up to translation (a condition arising from the dynamical Manin--Mumford conjecture), we have $\mathrm{ed}(f^s)=\mathrm{dim}(A)$ for some integer $s>0$. Our examples also show the necessity of both the hypothesis and iteration. We also give an affirmative answer to Koll\'ar and Zhuang's original question when $A$ is a simple abelian surface and $f$ is not $2$-polarized.

math.AG

InnoGym: Benchmarking the Innovation Potential of AI Agents

LLMs and Agents have achieved impressive progress in code generation, mathematical reasoning, and scientific discovery. However, existing benchmarks primarily measure correctness, overlooking the diversity of methods behind solutions. True innovation depends not only on producing correct answers but also on the originality of the approach. We present InnoGym, the first benchmark and framework designed to systematically evaluate the innovation potential of AI agents. InnoGym introduces two complementary metrics: performance gain, which measures improvement over the best-known solutions, and novelty, which captures methodological differences from prior approaches. The benchmark includes 18 carefully curated tasks from real-world engineering and scientific domains, each standardized through resource filtering, evaluator validation, and solution collection. In addition, we provide iGym, a unified execution environment for reproducible and long-horizon evaluations. Extensive experiments show that while some agents produce novel approaches, their lack of robustness limits performance gains. These results highlight a key gap between creativity and effectiveness, underscoring the need for benchmarks that evaluate both.

cs.CL

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations

Replicating AI research is a crucial yet challenging task for large language model (LLM) agents. Existing approaches often struggle to generate executable code, primarily due to insufficient background knowledge and the limitations of retrieval-augmented generation (RAG) methods, which fail to capture latent technical details hidden in referenced papers. Furthermore, previous approaches tend to overlook valuable implementation-level code signals and lack structured knowledge representations that support multi-granular retrieval and reuse. To overcome these challenges, we propose Executable Knowledge Graphs (xKG), a pluggable, paper-centric knowledge base that automatically integrates code snippets and technical insights extracted from scientific literature. When integrated into three agent frameworks with two different LLMs, xKG shows substantial performance gains (10.9% with o3-mini) on PaperBench, demonstrating its effectiveness as a general and extensible solution for automated AI research replication. Code is available at https://github.com/zjunlp/xKG.

cs.CL

Interactive Recommendation Agent with Active User Commands

Traditional recommender systems rely on passive feedback mechanisms that limit users to simple choices such as like and dislike. However, these coarse-grained signals fail to capture users' nuanced behavior motivations and intentions. In turn, current systems cannot also distinguish which specific item attributes drive user satisfaction or dissatisfaction, resulting in inaccurate preference modeling. These fundamental limitations create a persistent gap between user intentions and system interpretations, ultimately undermining user satisfaction and harming system effectiveness. To address these limitations, we introduce the Interactive Recommendation Feed (IRF), a pioneering paradigm that enables natural language commands within mainstream recommendation feeds. Unlike traditional systems that confine users to passive implicit behavioral influence, IRF empowers active explicit control over recommendation policies through real-time linguistic commands. To support this paradigm, we develop RecBot, a dual-agent architecture where a Parser Agent transforms linguistic expressions into structured preferences and a Planner Agent dynamically orchestrates adaptive tool chains for on-the-fly policy adjustment. To enable practical deployment, we employ simulation-augmented knowledge distillation to achieve efficient performance while maintaining strong reasoning capabilities. Through extensive offline and long-term online experiments, RecBot shows significant improvements in both user satisfaction and business outcomes.

cs.IR

RecGPT Technical Report

Recommender systems are among the most impactful applications of artificial intelligence, serving as critical infrastructure connecting users, merchants, and platforms. However, most current industrial systems remain heavily reliant on historical co-occurrence patterns and log-fitting objectives, i.e., optimizing for past user interactions without explicitly modeling user intent. This log-fitting approach often leads to overfitting to narrow historical preferences, failing to capture users' evolving and latent interests. As a result, it reinforces filter bubbles and long-tail phenomena, ultimately harming user experience and threatening the sustainability of the whole recommendation ecosystem. To address these challenges, we rethink the overall design paradigm of recommender systems and propose RecGPT, a next-generation framework that places user intent at the center of the recommendation pipeline. By integrating large language models (LLMs) into key stages of user interest mining, item retrieval, and explanation generation, RecGPT transforms log-fitting recommendation into an intent-centric process. To effectively align general-purpose LLMs to the above domain-specific recommendation tasks at scale, RecGPT incorporates a multi-stage training paradigm, which integrates reasoning-enhanced pre-alignment and self-training evolution, guided by a Human-LLM cooperative judge system. Currently, RecGPT has been fully deployed on the Taobao App. Online experiments demonstrate that RecGPT achieves consistent performance gains across stakeholders: users benefit from increased content diversity and satisfaction, merchants and the platform gain greater exposure and conversions. These comprehensive improvement results across all stakeholders validates that LLM-driven, intent-centric design can foster a more sustainable and mutually beneficial recommendation ecosystem.

cs.IR