Searcharxiv⌕ Search

arXiv subjects

Xiaoyuan Zhang

Publications and source records attributed to Xiaoyuan Zhang.

At least 19 recordsLinked to original sources

What Matters in Designing World Action Models: An Empirical Study

World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual contributions and systematically compare alternative designs. In this work, we present a controlled study that disentangles these design choices and analyzes not only their empirical effects, but also how and why they shape WAMs. More specifically, we focus on three fundamental questions in building WAMs: (1) what causal structure should govern the interaction between world modeling and action generation? (2) in which latent space should world modeling be performed? and (3) how do different world-action modeling objectives affect model behavior and performance? Through structurally controlled experiments on three representative benchmarks, RoboCasa-GR1, LIBERO, and LIBERO-Plus, we systematically compare six causal structures, eight latent representations, and four training objectives, covering popular design choices in existing WAMs. We further validate our key findings on real-robot data from the DROID dataset. We hope to provide a systematic understanding of how core design choices affect world-action modeling and what principles can guide the development of future WAM systems.

cs.RO↗

From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity

The ambitious requirements of sixth-generation (6G) networks are driving communication systems from reliable bit delivery toward meaning-aware and task-oriented connectivity. Large models (LMs), with strong multimodal understanding and generation capabilities, have accelerated this shift and made semantic communication (SemCom) increasingly practical. Yet current LM-driven SemCom remains fragmented: semantic representations are typically tied to specific modalities, models, or tasks. While the bit provides a universal unit for digital transport, there is still no analogous unit for representing and processing semantics, which limits interoperability, theoretical unification, and scalable system design. We argue that tokens provide a natural candidate for this missing abstraction. Two trends support this: unified multimodal LMs now encode text, images, audio, video, and robot actions in one token space, while distributed LM inference already generates substantial token-level traffic through expert routing, cache transfer, and speculative decoding. Token communication (TokenCom) emerges by unifying these trends, using the LM's native processing unit as a communication abstraction above the bit level and enabling importance assignment, error handling, and resource allocation directly at token granularity. This survey traces the evolution from LM-driven SemCom to TokenCom. We review three major directions of LM-driven SemCom: source-centric semantic coding, channel semantics for physical-layer tasks, and collaborative edge-device intelligence. We then examine the token abstraction, the transmission techniques it requires, and two emerging paradigms, namely TokenCom for LM services and for embodied and agentic intelligence. Finally, we identify open challenges toward unified, scalable, and AI-native 6G communication systems.

eess.SP↗

MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning

Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured combinatorial spaces. Traditional evolutionary algorithms often get trapped in local optima, while expert knowledge can provide crucial guidance for accelerating convergence. Large language models (LLMs) offer powerful priors and reasoning ability, making them natural optimizers when expert knowledge matters. However, closed-source LLMs, though strong in exploration, cannot update their parameters and thus cannot internalize experience. Conversely, smaller open models can be continually fine-tuned but lack broad knowledge and reasoning strength. We introduce Multi-LLM Collaborative Co-evolution (MCCE), a hybrid framework that unites a frozen closed-source LLM with a lightweight trainable model. The system maintains a trajectory memory of past search processes; the small model is progressively refined via reinforcement learning, with the two models jointly supporting and complementing each other in global exploration. Unlike model distillation, this process enhances the capabilities of both models through mutual inspiration. Experiments on multi-objective drug design benchmarks show that MCCE achieves state-of-the-art Pareto front quality and consistently outperforms baselines. These results highlight a new paradigm for enabling continual evolution in hybrid LLM systems, combining knowledge-driven exploration with experience-driven learning.

cs.LG↗

Grounded Normative Rule Generation with Structured Search

Normative rules like institutional charters and workplace policies must be both human-readable and operationally verifiable against actual environment records. However, current language generation and structured-output benchmarks primarily reward surface fluency or schema compliance, leaving operational grounding weakly tested. This creates a critical vulnerability where standard language models generate plausible-sounding policies that fail during enforcement because they rely on unavailable data logs or misaligned scopes. To address this challenge, we formalize the problem as Grounded Normative Rule Synthesis (GNRS) and introduce GNRS-Search, a framework that utilizes Markov Chain Monte Carlo (MCMC) sampling to optimize a discrete, five-slot And-Or Graph (AOG). By explicitly decoupling intermediate operational structure from final prose generation, this method isolates executable feasibility from writing style and allows rule failures to be localized prior to surface realization. We evaluate our approach on GNRS-Bench, a benchmark spanning 116 controlled goals across eight scene families, and RealCharter-Bench, which evaluates transfer to 53 real-derived policy tasks with hidden source clauses. GNRS-Search raises average rubric quality from 68.8% to 81.0% and ranks first under a disclosed executable composite metric, while systematic slot interventions confirm that performance gains stem from robust operational logic rather than rhetorical tuning. Ultimately, by transforming automated rule drafting into an inspectable search problem, this work provides a foundational paradigm for deploying verifiable and compliance-ready personal agents within regulated environments.

cs.CL↗

Energy Correlators in $V + X$ as a Benchmark Observable for Precision QCD

We propose projected energy correlators measured on the recoiling QCD radiation of a $Z/γ$ as a benchmark observable for precision QCD at the LHC. Using the $Z/γ$ as a hard scale prevents the need for a jet algorithm, simplifying both the perturbative and non-perturbative corrections. We develop a framework to combine state-of-the-art fixed-order amplitudes, high order resummation, and universal non-perturbative matrix elements. Our approach is based on numerically computed inclusive hard functions, allowing flexibility in the process and the inclusion of realistic experimental cuts. We perform detailed numerical studies of the projected energy correlators at next-to-leading order + next-to-next-to-leading logarithm (NLO+NNLL) to verify the stability of our setup. We present numerical results at NLO+NNLL, which are the first complete matched predictions for energy correlators at the LHC at this order. We discuss the prospects for extensions to higher orders, outlining a path to NNLO calculations of energy correlators at the LHC.

hep-ph↗

Heavy Jet Mass in Hadronic Higgs Decays

The heavy jet mass distributions in hadronic Higgs decays are computed to next-to-next-to-next-to-leading logarithmic order (N${}^3$LL${}^\prime$) in the dijet limit and NNLL in the trijet limit, matched to the next-to-next-to-leading order (NNLO). Both resummation results are obtained from the factorization theorems in Soft-Collinear Effective Theory. In particular, we study the Sudakov shoulders in the trijet region, originating from the incomplete cancellation of infrared singularities between final states with different parton multiplicities, and resum the induced large logarithms to all orders. The shoulder resummation yields sizable corrections and improves the perturbative stability of the distribution. Our results provide state-of-the-art predictions for heavy jet mass in $H \to gg$ and $H\to q\bar{q}$ and can be applied to precision Higgs measurements at future $e^+e^-$ colliders.

hep-ph↗

PILA: Plug-and-Play Insertion for LLM-native Advertising

How to monetize large language models (LLMs) by naturally integrating sponsored content into their responses, known as LLM-native advertising, has recently emerged as a critical problem. However, existing solutions entangle advertising with content generation inside a single model, which is incompatible with modern API-only or workflow-based LLM applications and inevitably compromises the original response quality. To address this, we propose PILA, which reformulates ad insertion as a conditional response rewriting problem and decouples it from the upstream service as a lightweight sidecar module. PILA is model-agnostic and can be seamlessly integrated with existing LLM services without modifying the base model or its workflow. It also exposes a controllable trade-off between user-side naturalness and ad-side exposure, offering a practical interface for downstream pricing and deployment. Experiments across diverse upstream models show that \pila consistently improves ad effectiveness while preserving response quality, highlighting its promise as a practical solution for LLM-native advertising.

cs.CL↗

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining

Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization geometry, state dynamics, numerical stability, implementation constraints, and empirical generalization. Existing automated optimizer discovery methods typically search either over unconstrained code spaces or within narrowly parameterized optimizer families. The former is flexible but often produces invalid or uninterpretable programs, while the latter is stable but limits novelty. We introduce OPTScientist, a theory-guided multi-agent framework for optimizer discovery in a typed domain-specific language (DSL). OPTScientist formulates optimizer design as a constrained scientific search process, where candidate updates are expressed through direction, scaling, preconditioning, regularization, state, and grouping modules. Four role agents, Theorist, Designer, Engineer, and Reviewer, collaborate within a single orchestration loop to propose hypotheses, synthesize DSL candidates, compile and evaluate optimizers, and critique results. To overcome the limitations of a fixed search space, OPTScientist combines evolutionary search over optimizer programs with a second-stage mechanism that proposes small DSL extensions when repeated failures reveal representational bottlenecks. Using this framework, we discover RS-MR, a reduced-state matrix optimizer that improves transformer pretraining over strong baselines under our native evaluation protocol. Our results suggest a path toward automated optimizer science grounded in theory, typed programs, compiler validation, and closed-loop experimentation.

cs.AI↗

WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error refinement. A natural way to improve search efficiency is to use world modeling, which can help identify promising optimization directions before costly evaluation. Large language models can predict the outcomes of these candidates with nontrivial accuracy because of their implicit knowledge. Motivated by this observation, we propose WMLLM, a self-evolving optimization-agent framework based on predict-then-act world modeling. The agent first predicts promising directions and then acts to generate candidates. Combined with agentic multi-turn refinement, population-based search, and reinforcement learning, WMLLM refines both its implicit world model and its optimization strategy during search. Experiments on black-box optimization tasks, especially multi-objective molecular optimization, show that WMLLM improves sample efficiency and final optimization performance. On the multi-objective molecular optimization benchmark, WMLLM achieves state-of-the-art results under a limited evaluation budget.

cs.LG↗

Projected Energy Correlators: Two-Loop Jet Functions and NNLL Resummation

We present the next-to-next-to-leading logarithmic (NNLL) collinear resummation of projected $N$-point energy correlators (ENCs) up to $N=6$, matched to fixed-order predictions at NLO, in both electron-positron annihilation and Higgs decay to gluons. The key new ingredient is the two-loop jet function for $N=4,5,6$, which we compute semi-analytically using Integration-by-Parts and differential equations. We further include the leading non-perturbative corrections for ENCs, described by two universal soft matrix elements $\overlineΩ_{1q},\overlineΩ_{1g}$ of order $Λ_{\rm QCD}$, whose evolution is governed by anomalous dimensions for $(N-1)$-point correlators. The matched distributions are compared with parton-shower simulations from Pythia8 and Herwig7, and we study the sensitivity of both the absolute spectra and their ratios to the two-point energy correlator under variations of $α_s$ and $\overlineΩ_{1q,1g}$. Our results show that higher-point projected energy correlators are now under quantitative control at NNLL accuracy, opening the door to future $α_s$ extractions with complementary systematics.

hep-ph↗

Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems

Model-based planning in robotic domains is challenged by the hybrid nature of physical dynamics, where continuous motion is punctuated by discrete events such as contacts and impacts. Conventional latent world models typically employ monolithic neural networks that enforce global continuity, which over-smooths distinct dynamic modes (e.g., sticking vs. sliding, flight vs. stance). For a planner, this smoothing results in compounding errors during long-horizon lookaheads, rendering the search process unreliable at physical boundaries. To address this, we introduce the Prismatic World Model (PRISM-WM), a structured architecture designed to decompose complex hybrid dynamics into composable primitives. PRISM-WM uses a context-aware Mixture-of-Experts (MoE) framework where a gating mechanism implicitly identifies the current physical mode, and specialized experts predict the associated transition dynamics. We further introduce a latent orthogonalization objective to ensure expert diversity, preventing mode collapse. By modeling the mode transitions in system dynamics, PRISM-WM reduces rollout drift. Experiments on continuous control benchmarks, including high-dimensional humanoids and multi-task settings, demonstrate that PRISM-WM provides a high-fidelity substrate for trajectory optimization algorithms (e.g., TD-MPC), indicating its potential as a foundational model for model-based agents.

cs.AI↗

MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers

Machine learning methods have shown that data-driven policies can accelerate mixed-integer linear programming (MILP) solvers, but many such approaches remain difficult to inspect, adapt, and deploy because the learned policy is represented as an external predictor or other opaque model. By contrast, explicit solver logic is easier to understand and integrate, but is usually hand-designed rather than learned from solver feedback. We study whether the automatic design of MILP solver logic can instead be cast as LLM-guided closed-loop search over executable white-box components evaluated directly by end-to-end solver behavior. To this end, we propose a closed-loop program evolution framework for MILP solver auto-design, implemented through PySCIPOpt, and instantiate it on the joint design of a cut selector and a branching rule. Candidate programs are iteratively generated, loaded into SCIP, and evaluated by direct execution on MILP instances, with the resulting feedback guiding performance-based selection, targeted repair, diagnostic reflection, and diversity-aware population maintenance. The method outputs explicit solver components that can be inspected, modified, and deployed within standard solver workflows. Across four benchmark families, we find that LLM-guided program evolution can discover competitive domain-specialized policies in several settings.

cs.AI↗

A Precise Determination of $α_s$ from the Heavy Jet Mass Distribution

A global fit for $α_s(m_Z)$ is performed on available $e^+e^-$ data for the heavy jet mass distribution. The state-of-the-art theory prediction includes $\mathcal{O}(α_s^3)$ fixed-order results, N$^3$LL$^\prime$ dijet resummation, N$^2$LL Sudakov shoulder resummation, and a first-principles treatment of power corrections in the dijet region. Theoretical correlations are incorporated through a flat random-scan covariance matrix. The global fit results in $0.1148^{+ 0.0015}_{-0.0022}$, compatible with similar determinations from thrust and $C$-parameter. Dijet resummation is essential for a robust fit, as it engenders insensitivity to the fit-range lower cutoff; without resummation the fit-range sensitivity is overwhelming. In addition, we find evidence for a negative power correction in the trijet region if and only if Sudakov shoulder resummation is included.

hep-ph↗

Les Houches study on inclusive jet production at NNLO+NNLL

Jet production at the LHC is a powerful probe of QCD, making it ideal for precision tests and determinations of QCD parameters such as parton distribution functions and the strong coupling constant. To make the most of the abundant jet production data collected at the LHC, precise calculations are required. While state-of-the-art calculations reach next-to-next-to-leading order (NNLO) QCD accuracy, a critical assessment of the remaining uncertainties arising from non-perturbative effects and missing higher orders remains crucial for correctly interpreting comparisons between theory and data. Scale variation is nearly always used to determine effects from missing higher orders. In this article, we reassess this method in the context of inclusive jet production by performing NNLO QCD calculations supplemented by small-jet-radius resummation through next-to-next-to-leading-logarithmic accuracy (NNLL). We find that NNLL resummation can have an appreciable impact on the scale uncertainty for inclusive jet cross sections, and, for some scale choices, can lead to sizeable shifts of the central cross section. We conclude that scale variations in fixed-order and resummed calculations can drastically underestimate the impact of higher orders for commonly used jet radius parameters, and that missing higher-order estimates obtained via scale variations should be considered unreliable. Our findings add further evidence to the importance of going beyond scale variations in jet and jet substructure calculations.

hep-ph↗

Aegis: Automated Error Generation and Attribution for Multi-Agent Systems

Large language model based multi-agent systems (MAS) have unlocked significant advancements in tackling complex problems, but their increasing capability introduces a structural fragility that makes them difficult to debug. A key obstacle to improving their reliability is the severe scarcity of large-scale, diverse datasets for error attribution, as existing resources rely on costly and unscalable manual annotation. To address this bottleneck, we introduce Aegis, a novel framework for Automated error generation and attribution for multi-agent systems. Aegis constructs a large dataset of 9,533 trajectories with annotated faulty agents and error modes, covering diverse MAS architectures and task domains. This is achieved using a LLM-based manipulator that can adaptively inject context-aware errors into successful execution trajectories. Leveraging fine-grained labels and the structured arrangement of positive-negative sample pairs, Aegis supports three different learning paradigms: Supervised Fine-Tuning, Reinforcement Learning, and Contrastive Learning. We develop learning methods for each paradigm. Comprehensive experiments show that trained models consistently achieve substantial improvements in error attribution. Notably, several of our fine-tuned LLMs demonstrate performance competitive with or superior to proprietary models an order of magnitude larger, validating our automated data generation framework as a crucial resource for developing more robust and interpretable multi-agent systems. Our project website is available at https://kfq20.github.io/Aegis-Website/.

cs.RO↗

Adaptive Online Mirror Descent for Tchebycheff Scalarization in Multi-Objective Learning

Multi-objective learning (MOL) aims to learn under multiple potentially conflicting objectives and strike a proper balance. While recent preference-guided MOL methods often rely on additional optimization objectives or constraints, we consider the classic Tchebycheff scalarization (TCH) that naturally allows for locating solutions with user-specified trade-offs. Due to its minimax formulation, directly optimizing TCH often leads to training oscillation and stagnation. In light of this limitation, we propose an adaptive online mirror descent algorithm for TCH, called (Ada)OMD-TCH. One of our main ingredients is an adaptive online-to-batch conversion that significantly improves solution optimality over traditional conversion in practice while maintaining the same theoretical convergence guarantees. We show that (Ada)OMD-TCH achieves a convergence rate of $\mathcal O(\sqrt{\log m/T})$, where $m$ is the number of objectives and $T$ is the number of rounds, providing a tighter dependency on $m$ in the offline setting compared to existing work. Empirically, we demonstrate on both synthetic problems and federated learning tasks that (Ada)OMD-TCH effectively smooths the training process and yields preference-guided, specific, diverse, and fair solutions.

cs.LG↗

Precision Jet Substructure of Boosted Boson Decays with Energy Correlators

We initiate the precision study of boosted jet substructure using energy correlators, applying this framework to hadronic Higgs decays. We demonstrate that the two-body decay of the Higgs manifests as a distinct angular peak at $θ\sim \arccos(1-2/γ^2)$ for Lorentz boost factor $γ$. We show that infrared scales, such as the dead-cone effect and confinement transition, are also resolved within the boosted distribution. Precision analytic studies of boosted jet substructure may enable precision electroweak studies and open new avenues for new physics searches.

hep-ph↗

The life of central radio galaxies in clusters: AGN-ICM studies of eRASS1 clusters in the ASKAP fields

The mechanical feedback from the central AGNs can be crucial for balancing the radiative cooling of the intracluster medium at the cluster centre. We aim to understand the relationship between the power of AGN feedback and the cooling of gas in the centres of galaxy clusters by correlating the radio properties of the brightest cluster galaxies (BCGs) with the X-ray properties of their host clusters. We used catalogues from the first SRG/eROSITA All-Sky Survey (eRASS1) along with ASKAP radio data. In total, we identified 134 radio sources associated with BCGs of the 151 eRASS1 clusters located in the PS1, PS2, and SWAG-X ASKAP fields. Non-detections were treated as upper limits. We correlated BCG radio luminosity, largest linear size (LLS), and BCG offset with the integrated X-ray luminosity of their host clusters. To characterise cool cores (CCs) and non-cool cores (NCCs), we used the concentration parameter $c_{R_{500}}$ and combined it with the BCG offset to assess cluster dynamical state. We analysed the correlation between radio mechanical power and X-ray luminosity within the CC subsample. We observe a potential positive trend between LLS and BCG offset, suggesting an environmental effect on radio-source morphology. We find a weak trend where more luminous central radio galaxies are found in clusters with higher X-ray luminosity. Within the CC subsample, there is a positive but highly scattered relationship between the mechanical luminosity of AGN jets and the X-ray cooling luminosity. This finding is supported by bootstrap resampling and flux-flux analyses. The correlation indicates that AGN feedback is ineffective in high-luminosity (high-mass) clusters. At a cooling luminosity of $L_{\mathrm{X},~r<R_\mathrm{cool}}\approx 5.50\times10^{43}$ erg/s, on average, AGN feedback appears to contribute only about 13%-22% of the energy needed to offset the radiative losses in the ICM.

astro-ph.CO↗