Searcharxiv⌕ Search

arXiv subjects

Zijie Zhuang

Publications and source records attributed to Zijie Zhuang.

At least 19 recordsLinked to original sources

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.

cs.AI↗

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated artifacts may be unsupported or incomplete, executed rounds may be invalid or confounded, and later modifications may obscure earlier findings. We study \textbf{trajectory-to-evidence conversion}, asking what a completed research process has actually established. We introduce an evidence-grounded framework that couples bounded verification of consequential artifacts with post-execution claim qualification. A context-isolated generate--verify--repair process checks artifacts for evidence violations and missing downstream requirements before release. After execution, validity and attribution checks consolidate evidence across rounds, qualify intervention-level claims as actionable repairs, diagnostic guards, or withheld findings, and preserve admitted claims as auditable records with explicit provenance and applicability boundaries. A hybrid LLM-assisted controller subsequently applies, defers, or rejects records based on available target evidence. Record audits characterize which claims survive qualification, while downstream diagnostics identify affirmative applicability judgment as a bottleneck for the tested controller. Across paper-to-target adaptations, later rounds often improve on the first, while final rounds frequently underperform an earlier best, exposing non-monotonic trajectory evolution. Candidates produced through the complete workflow also yielded positive online lifts relative to deployed baselines.

cs.IR↗

RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation

Multimodal large language models (MLLMs) can convert multimodal item content into structured descriptions used as semantic features for recommendation. Conventional content-only generation, however, cannot use downstream user signals to determine which semantics should be emphasized. Recent user-conditioned methods incorporate these signals through user histories or profiles, but they require user information at inference and make generation user-dependent. In this paper, we introduce RecoReward, which instead uses behavior-derived rewards during training and preserves content-only inference. To instantiate this idea in live-stream recommendation, we treat historically engaged users as a proxy for future target users and use observational non-target users to estimate affinity shared broadly across users. The Recommender Affinity Score (RAS) contrasts these signals to provide user-selective feedback for reinforcement learning, allowing the learned policy to generate a single shared description without user inputs. In our offline benchmark, RecoReward-9B outperforms its Qwen3.5-9B baseline and all other evaluated models across seven recall metrics. Online A/B testing also shows performance gains. These results show that RecoReward trains the MLLM to produce item features that benefit downstream recommendation while retaining content-only serving.

cs.IR↗

Tightness of exponential metrics for log-correlated Gaussian fields in arbitrary dimension

We prove the tightness of a natural approximation scheme for an analog of the Liouville quantum gravity metric on $\mathbb R^d$ for arbitrary $d\geq 2$. More precisely, let $\{h_n\}_{n\geq 1}$ be a suitable sequence of Gaussian random functions which approximates a log-correlated Gaussian field on $\mathbb R^d$. Consider the family of random metrics on $\mathbb R^d$ obtained by weighting the lengths of paths by $e^{ξh_n}$, where $ξ> 0$ is a parameter. We prove that if $ξ$ belongs to the subcritical phase (which is defined by the condition that the distance exponent $Q(ξ)$ is greater than $\sqrt{2d}$), then after appropriate re-scaling, these metrics are tight and that every subsequential limit is a metric on $\mathbb R^d$ which induces the Euclidean topology. We include a substantial list of open problems.

math.PR↗

The bulk one-arm exponent for the CLE$_{κ'}$ percolations

The conformal loop ensemble (CLE) is a conformally invariant random collection of loops. In the non-simple regime $κ'\in (4,8)$, it describes the scaling limit of the critical Fortuin-Kasteleyn (FK) percolations. CLE percolations were introduced by Miller-Sheffield-Werner (2017). The CLE$_{κ'}$ percolations describe the scaling limit of a natural variant of the FK percolation called the fuzzy Potts model, which has an additional percolation parameter $r$. Based on CLE percolations and assuming that the convergence of the FK percolation to CLE, K{ö}hler-Schindler and Lehmkuehler (2022) derived all the arm exponents for the fuzzy Potts model except the bulk one-arm exponent. In this paper, we exactly solve this exponent, which prescribes the dimension of the clusters in CLE$_{κ'}$ percolations. As a special case, the bichromatic one-arm exponent for the critical 3-state Potts model should be $4/135$. To the best of our knowledge, this natural exponent was not predicted in physics. Our derivation relies on the iterative construction of CLE percolations from the boundary conformal loop ensemble (BCLE), and the coupling between Liouville quantum gravity and SLE curves. The source of the exact solvability comes from the structure constants of boundary Liouville conformal field theory. A key technical step is to prove a conformal welding result for the target-invariant radial SLE curves. As intermediate steps in our derivation, we obtain several exact results for BCLE in both the simple and non-simple regimes, which extend results of Ang-Sun-Yu-Zhuang (2024) on the touching probability of non-simple CLE. This also provides an alternative derivation of the relation between the BCLE parameter $ρ$ and the additional percolation parameter $r$ in CLE percolations, which was originally due to Miller-Sheffield-Werner (2021, 2022).

math.PR↗

AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly with headcount rather than compounding with evidence, compute, and accumulated experimental knowledge. We present AgentX, a production-deployed multi-agent system that fundamentally restructures this production function. AgentX operates as a self-evolving development engine: it autonomously generates, implements, evaluates, and learns from recommendation experiments at a scale and pace that no manual workflow can sustain. The system orchestrates four tightly coupled stages in a closed loop. A Brainstorm Agent synthesizes evidence from historical experiments, system architecture, data analysis, and external research into ranked, executable proposals. A Developing Agent translates each proposal into production-ready code through repository-grounded generation and multi-dimensional reliability verification. An Evaluation Agent conducts safe online rollout with guardrail-vetoed A/B judgment, converting both successes and failures into structured knowledge assets. A Harness Evolution layer (SGPO) then distills execution trajectories into semantic-gradient updates that continuously sharpen the agents themselves -- making the system not merely automated, but self-improving.

cs.AI↗

Foundation Protocol: A Coordination Layer for Agentic Society

Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly interact with one another. As these systems scale, the bottleneck shifts away from raw model capability toward coordination. Agents need to form reliable relationships, organize multi-agent work, exchange value, support an AI economy, and stay safe and accountable under real-world oversight. This paper introduces the Foundation Protocol (FP), a graph-first coordination layer for an emerging human-AI society. FP unifies heterogeneous entities, including agents, tools, resources, humans, institutions, and organizations, and supports native multi-party organization and event-based collaboration. It also provides economic primitives for metering, receipts, and settlement, and treats policy, provenance, and audit as first-class concerns. FP is designed to wrap and bridge existing protocols rather than replace them, enabling incremental adoption while reducing integration and governance overhead. The aim is to keep autonomous agency composable while keeping accountability non-negotiable, so that coordination itself can become shared infrastructure for a human-AI society that is open, pluralistic, and governable.

cs.AI↗

AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines

The performance of autonomous Web GUI agents heavily relies on the quality and quantity of their training data. However, a fundamental bottleneck persists: collecting interaction trajectories from real-world websites is expensive and difficult to verify. The underlying state transitions are hidden, leading to reliance on inconsistent and costly external verifiers to evaluate step-level correctness. To address this, we propose AutoWebWorld, a novel framework for synthesizing controllable and verifiable web environments by modeling them as Finite State Machines (FSMs) and use coding agents to translate FSMs into interactive websites. Unlike real websites, where state transitions are implicit, AutoWebWorld explicitly defines all states, actions, and transition rules. This enables programmatic verification: action correctness is checked against predefined rules, and task success is confirmed by reaching a goal state in the FSM graph. AutoWebWorld enables a fully automated search-and-verify pipeline, generating over 11,663 verified trajectories from 29 diverse web environments at only $0.04 per trajectory. Training on this synthetic data significantly boosts real-world performance. Our 7B Web GUI agent outperforms all baselines within 15 steps on WebVoyager. Furthermore, we observe a clear scaling law: as the synthetic data volume increases, performance on WebVoyager and Online-Mind2Web consistently improves.

cs.AI↗

SARM: LLM-Augmented Semantic Anchor for End-to-End Live-Streaming Ranking

Large-scale live-streaming recommendation requires precise modeling of non-stationary content semantics under strict real-time serving constraints. In industrial deployment, two common approaches exhibit fundamental limitations: discrete semantic abstractions sacrifice descriptive precision through clustering, while dense multimodal embeddings are extracted independently and remain weakly aligned with ranking optimization, limiting fine-grained content-aware ranking. To address these limitations, we propose \textbf{SARM}, an end-to-end ranking architecture that integrates natural-language semantic anchors directly into ranking optimization, enabling fine-grained author representations conditioned on multimodal content. Each semantic anchor is represented as learnable text tokens jointly optimized with ranking features, allowing the model to adapt content descriptions to ranking objectives. A lightweight dual-token gated design captures domain-specific live-streaming semantics, while an asymmetric deployment strategy preserves low-latency online training and serving. Extensive offline evaluation and large-scale A/B tests show consistent improvements over production baselines. SARM is fully deployed and serves over 400 million users daily.

cs.IR↗

Percolation of thick points of the log-correlated Gaussian field in high dimensions

We prove that the set of thick points of the log-correlated Gaussian field contains an unbounded path in sufficiently high dimensions. This contrasts with the two-dimensional case, where Aru, Papon, and Powell (2023) showed that the set of thick points is totally disconnected. This result has an interesting implication for the exponential metric of the log-correlated Gaussian field: in sufficiently high dimensions, when the parameter $ξ$ is large, the set-to-set distance exponent (if it exists) is negative. This suggests that a new phase may emerge for the exponential metric, which does not appear in two dimensions. In addition, we establish similar results for the set of thick points of the branching random walk. As an intermediate result, we also prove that the critical probability for fractal percolation converges to 0 as $d \to \infty$.

math.PR↗

VideoMemory: Toward Consistent Video Generation via Memory Integration

Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high-quality short clips but often fail to preserve entity identity and appearance when scenes change or when entities reappear after long temporal gaps. We present VideoMemory, an entity-centric framework that integrates narrative planning with visual generation through a Dynamic Memory Bank. Given a structured script, a multi-agent system decomposes the narrative into shots, retrieves entity representations from memory, and synthesizes keyframes and videos conditioned on these retrieved states. The Dynamic Memory Bank stores explicit visual and semantic descriptors for characters, props, and backgrounds, and is updated after each shot to reflect story-driven changes while preserving identity. This retrieval-update mechanism enables consistent portrayal of entities across distant shots and supports coherent long-form generation. To evaluate this setting, we construct a 54-case multi-shot consistency benchmark covering character-, prop-, and background-persistent scenarios. Extensive experiments show that VideoMemory achieves strong entity-level coherence and high perceptual quality across diverse narrative sequences.

cs.CV↗

Weak exponential metrics for high-dimensional log-correlated Gaussian fields

For log-correlated Gaussian fields on $\mathbb{R}^d$ with $d \geq 2$, Ding-Gwynne-Zhuang (2023) established the existence of subsequential limits of exponential metrics obtained from appropriate approximations. For $γ\in (0,\sqrt{2d})$, we define a \textit{weak $γ$-exponential metric} to be a map $h \mapsto D_h$ that assigns to a sample of a log-correlated Gaussian field $h$ a continuous metric on $\mathbb{R}^d$ satisfying a list of axioms. We prove that every subsequential limit of exponential metrics built from appropriate approximations of $h$ is a weak $γ$-exponential metric in this sense. Moreover, we establish general properties that hold for any weak exponential metric: (1). sharp moment bounds for several natural distances; (2). optimal Hölder exponents when comparing $D_h$ and the Euclidean metric; and (3). Hausdorff dimension and a KPZ relation. These results extend the two-dimensional Liouville quantum gravity metric theory to higher dimensions. Along the way we derive several useful properties for log-correlated Gaussian fields including the equivalence between white-noise decomposition and convolution, and a shell independence lemma.

math.PR↗

Dynamical random field Ising model at zero temperature

In this paper, we study the evolution of the zero-temperature random field Ising model as the mean of the external field $M$ increases from $-\infty$ to $\infty$. We focus on two types of evolutions: the ground state evolution and the Glauber evolution. For the ground state evolution, we investigate the occurrence of global avalanche, a moment where a large fraction of spins flip simultaneously from minus to plus. In two dimensions, no global avalanche occurs, while in three or higher dimensions, there is a phase transition: a global avalanche happens when the noise intensity is small, but not when it is large. Additionally, we study the zero-temperature Glauber evolution, where spins are updated locally to minimize the Hamiltonian. Our results show that for small noise intensity, in dimensions $d =2$ or $3$, most spins flip around a critical time $c_d = \frac{2 \sqrt{d}}{1 + \sqrt{d}}$ (but we cannot decide whether such flipping occurs simultaneously or not). We also connect this process to polluted bootstrap percolation and solve an open problem on it.

math.PR↗

Schramm-Loewner evolution contains a topological Sierpiński carpet when $κ$ is close to 8

We consider the Schramm-Loewner evolution (SLE$_κ$) for $κ\in (4,8)$, which is the regime where the curve is self-intersecting but not space-filling. We show that there exists $δ_0>0$ such that for $κ\in (8 - δ_0,8)$, the range of an SLE$_κ$ curve almost surely contains a topological Sierpiński carpet. Combined with a result of Ntalampekos (2021), this implies that in this parameter range, SLE$_κ$ is almost surely conformally non-removable, and the conformal welding problem for SLE$_κ$ does not have a unique solution. Our result also implies that for $κ\in (8 - δ_0,8)$, the adjacency graph of the complementary connected components of the SLE$_κ$ curve is disconnected.

math.PR↗

Three-point connectivity constant for $q$-state Potts spin clusters

Recently, Ang--Cai--Sun--Wu (2024) determined the three-point connectivity constant for two-dimensional critical percolation, confirming a prediction of Delfino and Viti (2010). In this paper, we address the analogous problem for planar critical $q$-state Potts spin clusters. We introduce a continuum three-point connectivity constant and compute it explicitly. Under the scaling-limit conjecture for Potts spin clusters, this quantity coincides with the scaling limit of the properly normalized probability that three points lie in the same spin cluster. The resulting formula agrees with the imaginary DOZZ formula up to an explicit $q$-dependent constant with a geometric interpretation. This answers a question from Delfino--Picco--Santachiara--Viti (2013). The proof exploits the coupling between CLE and LQG, together with the BCLE descriptions of $q$-state Potts scaling limits due to Miller--Sheffield--Werner (2017) and Köhler-Schindler and Lehmkühler (2025).

math.PR↗

Bounds on the distance exponent for higher-dimensional Liouville first passage percolation

For $ξ\geq 0$ and $d \geq 3$, the higher-dimensional Liouville first passage percolation (LFPP) is a random metric on $ε\mathbb{Z}^d$ obtained by reweighting each vertex by $e^{ξh_ε(x)}$, where $h_ε(x)$ is a continuous mollification of the whole-space log-correlated Gaussian field. This metric generalizes the two-dimensional LFPP, which is related to Liouville quantum gravity. We derive several estimates for the set-to-set distance exponent of this metric, including upper and lower bounds and bounds on its derivative with respect to $ξ$. In the subcritical region for $ξ$, we derive estimates for the fractal dimension and show that it is continuous and strictly increasing with respect to $ξ$. In particular, our result is an important step towards proving a technical assumption made in previous work by the first author and Gwynne. These are also the first bounds on the distance exponent for LFPP in higher dimensions.

math.PR↗

Mixing rate exponent of planar Fortuin-Kasteleyn percolation

Duminil-Copin and Manolescu (2022) recently proved the scaling relations for planar Fortuin-Kasteleyn (FK) percolation. In particular, they showed that the one-arm exponent and the mixing rate exponent are sufficient to derive the other near-critical exponents. The scaling limit of critical FK percolation is conjectured to be a conformally invariant random collection of loops called the conformal loop ensemble (CLE). In this paper, we define the CLE analog of the mixing rate exponent. Assuming the convergence of FK percolation to CLE, we show that the mixing rate exponent for FK percolation agrees with that of CLE. We prove that the CLE$_κ$ mixing rate exponent equals $\frac{3 κ}{8}-1$, thereby answering Question 3 of Duminil-Copin and Manolescu (2022). The derivation of the CLE exponent is based on an exact formula for the Radon-Nikodym derivative between the marginal laws of the odd-level and even-level CLE loops, which is obtained from the coupling between Liouville quantum gravity and CLE.

math.PR↗

Annulus crossing formulae for critical planar percolation

We derive exact formulae for three basic annulus crossing events for the critical planar Bernoulli percolation in the continuum limit. The first is for the probability that there is an open path connecting the two boundaries of an annulus of inner radius $r$ and outer radius $R$. The second is for the probability that there are both open and closed paths connecting the two annulus boundaries. These two results were predicted by Cardy based on non-rigorous Coulomb gas arguments. Our third result gives the probability that there are two disjoint open paths connecting the two boundaries. Its leading asymptotic as $r/R\to 0$ is captured by the so-called backbone exponent, a transcendental number recently determined by Nolin, Qian and two of the authors. This exponent is the unique real root to the equation $\frac{\sqrt{36 x +3}}{4} + \sin (\frac{2 π\sqrt{12 x +1}}{3} ) =0$, other than $-\frac{1}{12}$ and $\frac{1}{4}$. Besides these three real roots, this equation has countably many complex roots. Our third result shows that these roots appear exactly as exponents of the subleading terms in the crossing formula. This suggests that the backbone exponent is part of a conformal field theory (CFT) whose bulk spectrum contains this set of roots. Expanding the same crossing probability as $r/R\to 1$, we obtain a series with logarithmic corrections at every order, suggesting that the backbone exponent is related to a logarithmic boundary CFT. Our proofs are based on the coupling between SLE curves and Liouville quantum gravity (LQG). The key is to encode the annulus crossing probabilities by the random moduli of certain LQG surfaces with annular topology, whose law can be extracted from the dependence of the LQG annuli partition function on their boundary lengths.

math.PR↗