Searcharxiv⌕ Search

arXiv subjects

Xiangrong Zhu

Publications and source records attributed to Xiangrong Zhu.

At least 19 recordsLinked to original sources

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents either rely on explicit user instructions or engage in multi-turn clarification to elicit user preferences. However, both approaches overlook the rich behavioral signals latent in users' past behaviors, which implicitly encode their preferences. This over-reliance on active user input increases interaction burden and limits plan personalization. To bridge this gap, we introduce a new task, Behavior-Aware Travel Planning, which infers user preferences directly from past behaviors and generates personalized travel plans. To facilitate research on this task, we introduce Behavior2Trip, a benchmark constructed from one of the largest Chinese online travel platforms, comprising 11,400 instances. Each instance represents an average of 39.8 past user behaviors spanning 14 attributes across 5 preference dimensions. We further propose B2T-Agent, a reinforcement learning-based agent that leverages user behavior trajectories, interacts with external tools for preference-aligned retrieval, and maintains an internal memory module. Experiments on Behavior2Trip show that GPT-4.1 achieves a full-constraint pass rate of only 0.5\% on the hardest tasks, while B2T-Agent built upon Qwen3-8B outperforms all baselines, highlighting the substantial challenge of this task. Moreover, Qwen3-8B trained with B2T-Agent also outperforms GPT-4.1 on the TravelPlanner benchmark, demonstrating strong generalization. Code and data are available at https://github.com/BUAA-IRIP-LLM/Behavior2Trip

cs.CL↗

Machine learning the impact parameter in heavy-ion collisions at $\sqrt{s_{\rm NN}}$ = 4 and 11 GeV: a cross-check study with UrQMD, AMPT, and JAM

By generating heavy-ion collision data with the ultrarelativistic quantum molecular dynamics (UrQMD) model, a multiphase transport (AMPT) model, and the JAM model, the impact parameter ($b$) in Au+Au collisions at $\sqrt{s_{\rm NN}}$ = 4 and 11 GeV is reconstructed using supervised learning and unsupervised learning in machine learning (ML). In supervised learning, the performance of ML algorithm is cross-checked by using data obtained from these three transport models. It is found that the typical mean absolute error (MAE) which measures the average magnitude of the absolute difference between the true and predicted $b$ is between 0.2-0.4 fm, even when training ML algorithm with data generated from one model but testing with data from others. While the conventional method (i.e., a polynomial fit to multiplicity as a function of $b$) only works for data generated from the same model. In the classification task, the present ML-based method also shows significantly superior results compared to the traditional approach. In unsupervised learning, the K-means clustering algorithm is used to partition collision events directly from experimental-style observables, showing that the algorithm autonomously identifies six clusters corresponding to different centrality classes without relying on predefined model-based binning. Our study demonstrates the strong robustness of using an ML algorithm trained on transport-model data for impact-parameter determination, and indicates that this method has the potential to be generalized to handle real experimental data.

nucl-th↗

Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representative methods such as GRPO assign uniform credit across all tokens, wasting gradient on routine tokens while under-crediting pivotal reasoning steps. Existing token-level credit assignment methods require resources beyond the model's own rollouts. GRPO variants rely on process reward models or ground-truth answers. Knowledge distillation assigns credit through per-token divergence but requires external teachers (On-Policy Distillation) or privileged information (On-Policy Self Distillation). However, these dependencies limit applicability in the pure RLVR setting. We observe that conditioning the model on its own verified trajectories induces a measurable per-token KL divergence between the original and conditioned distributions, and prove that distilling from a self-teacher constructed by verified trajectories leads to infeasible weighted-average solutions when multiple verified trajectories exist. We propose SC-GRPO (Self-Conditioned GRPO), which uses KL divergence mentioned before as a multiplicative weight on GRPO gradients. Across five benchmarks spanning math, code, and agentic tasks, SC-GRPO consistently outperforms 8.1% over GRPO and 5.9% over DAPO with stronger OOD performance. Moreover, SC-GRPO achieves higher performance than OPD.

cs.LG↗

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-quality training data. Existing approaches bootstrap from partial sources such as human-defined seeds or GitHub repositories to instantiate one component and then complete the rest, producing tasks confined to narrow seed distributions, environments misaligned with task semantics, and inefficient trajectories from unguided exploration. To address these limitations, we introduce Terminal-World, a fully automated pipeline that uses agent skills as the central synthesis primitive, which jointly encode what to accomplish, when to apply (preconditions and environment state), and how to execute, enabling task instructions, environments, and teacher trajectories to be co-derived. To further broaden the synthesis space, Terminal-World composes skills into skill teams and skill graphs for multi-role and cross-domain task synthesis. Using this pipeline, we construct 5,723 training environments and train Terminal-World-8B/14B/32B, evaluated across 6 benchmarks where the Terminal-World series consistently outperforms terminal-agent baselines. Notably, using the same teacher model and only 1.2% of the training data, Terminal-World-32B surpasses Nemotron-Terminal-32B on Terminal-Bench 2.0 by +4.5 Pass@1 (31.5) and achieves 43.8 Pass@3.

cs.CL↗

Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation

While large language model--powered agents can self-evolve by accumulating experience or by dynamically creating new assets (i.e., tools or expert agents), existing frameworks typically treat these two evolutionary processes in isolation. This separation overlooks their intrinsic interdependence: the former is inherently bounded by a manually predefined static toolset, while the latter generates new assets from scratch without experiential guidance, leading to limited capability growth and unstable evolution. To address this limitation, we introduce a novel paradigm of co-evolutionary Capability Expansion and Experience Distillation. Guided by this paradigm, we propose the \textbf{Mem$^{\textbf{2}}$Evolve}, which integrates two core components: \textbf{Experience Memory} and \textbf{Asset Memory}. Specifically, Mem$^{2}$Evolve leverages accumulated experience to guide the dynamic creation of assets, thereby expanding the agent's capability space while simultaneously acquiring new experience to achieve co-evolution. Extensive experiments across 6 task categories and 8 benchmarks demonstrate that Mem$^{2}$Evolve achieves improvement of 18.53\% over standard LLMs, 11.80\% over agents evolving solely through experience, and 6.46\% over those evolving solely through asset creation, establishing it as a substantially more effective and stable self-evolving agent framework. Code is available at: https://buaa-irip-llm.github.io/Mem2Evolve.

cs.CL↗

Linear Poisson Equations with Potential on Riemann Surfaces

We study interior estimates for solutions of the linear Poisson equation: $$ \triangle u = g u + f $$ where $g$ and $f$ belong to the Zygmund space $L\ln L$ on a Riemann surface $M$ satisfying the isoperimetric inequality. As applications, we derive corresponding interior estimates, Harnack inequalities, and a global estimate.

math.DG↗

Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition

In a rapidly evolving world where information updates swiftly, knowledge in large language models (LLMs) becomes outdated quickly. Retraining LLMs is not a cost-effective option, making knowledge editing (KE) without modifying parameters particularly necessary. We find that although existing retrieval-augmented generation (RAG)-based KE methods excel at editing simple knowledge, they struggle with KE in multi-hop question answering due to the issue of "edit skipping", which refers to skipping the relevant edited fact in inference. In addition to the diversity of natural language expressions of knowledge, edit skipping also arises from the mismatch between the granularity of LLMs in problem-solving and the facts in the edited memory. To address this issue, we propose a novel Iterative Retrieval-Augmented Knowledge Editing method with guided decomposition (IRAKE) through the guidance from single edited facts and entire edited cases. Experimental results demonstrate that IRAKE mitigates the failure of editing caused by edit skipping and outperforms state-of-the-art methods for KE in multi-hop question answering.

cs.CL↗

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models

Large Language Models (LLMs) have exhibited significant proficiency in code debugging, especially in automatic program repair, which may substantially reduce the time consumption of developers and enhance their efficiency. Significant advancements in debugging datasets have been made to promote the development of code debugging. However, these datasets primarily focus on assessing the LLM's function-level code repair capabilities, neglecting the more complex and realistic repository-level scenarios, which leads to an incomplete understanding of the LLM's challenges in repository-level debugging. While several repository-level datasets have been proposed, they often suffer from limitations such as limited diversity of tasks, languages, and error types. To mitigate this challenge, this paper introduces RepoDebug, a multi-task and multi-language repository-level code debugging dataset with 22 subtypes of errors that supports 8 commonly used programming languages and 3 debugging tasks. Furthermore, we conduct evaluation experiments on 10 LLMs, where Claude 3.5 Sonnect, the best-performing model, still cannot perform well in repository-level debugging.

cs.SE↗

RETAIL: Towards Real-world Travel Planning for Large Language Models

Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios. First, they assume users provide explicit queries, while in reality requirements are often implicit. Second, existing solutions ignore diverse environmental factors and user preferences, limiting the feasibility of plans. Third, systems can only generate plans with basic POI arrangements, failing to provide all-in-one plans with rich details. To mitigate these challenges, we construct a novel dataset \textbf{RETAIL}, which supports decision-making for implicit queries while covering explicit queries, both with and without revision needs. It also enables environmental awareness to ensure plan feasibility under real-world scenarios, while incorporating detailed POI information for all-in-one travel plans. Furthermore, we propose a topic-guided multi-agent framework, termed TGMA. Our experiments reveal that even the strongest existing model achieves merely a 1.0% pass rate, indicating real-world travel planning remains extremely challenging. In contrast, TGMA demonstrates substantially improved performance 2.72%, offering promising directions for real-world travel planning.

cs.AI↗

Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation

Counterfactual reasoning is crucial for robust video understanding but remains underexplored in existing multimodal benchmarks. In this paper, we introduce \textbf{COVER} (\textbf{\underline{CO}}unterfactual \textbf{\underline{V}}id\textbf{\underline{E}}o \textbf{\underline{R}}easoning), a multidimensional multimodal benchmark that systematically evaluates MLLMs across the abstract-concrete and perception-cognition dimensions. Beyond prior multimodal benchmarks, COVER decomposes complex queries into structured sub-questions, enabling fine-grained reasoning analysis. Experiments on commercial and open-source models reveal a strong correlation between sub-question accuracy and counterfactual reasoning performance, highlighting the role of structured inference in video understanding. Furthermore, our results suggest a key insight: enhancing the reasoning capability of models is essential for improving the robustness of video understanding. COVER establishes a new standard for assessing MLLMs' logical reasoning abilities in dynamic environments. Our work is available at https://github.com/gongyifan-hash/COVER-Benchmark.

cs.CV↗

Knowledge Graph-Guided Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising technology for addressing hallucination issues in the responses generated by large language models (LLMs). Existing studies on RAG primarily focus on applying semantic-based approaches to retrieve isolated relevant chunks, which ignore their intrinsic relationships. In this paper, we propose a novel Knowledge Graph-Guided Retrieval Augmented Generation (KG$^2$RAG) framework that utilizes knowledge graphs (KGs) to provide fact-level relationships between chunks, improving the diversity and coherence of the retrieved results. Specifically, after performing a semantic-based retrieval to provide seed chunks, KG$^2$RAG employs a KG-guided chunk expansion process and a KG-based chunk organization process to deliver relevant and important knowledge in well-organized paragraphs. Extensive experiments conducted on the HotpotQA dataset and its variants demonstrate the advantages of KG$^2$RAG compared to existing RAG-based approaches, in terms of both response quality and retrieval quality.

cs.CL↗

Endpoint regularity of general Fourier integral operators

Let $n\geq 1,0<ρ<1, \max\{ρ,1-ρ\}\leq δ\leq 1$ and $$m_1=ρ-n+(n-1)\min\{\frac 12,ρ\}+\frac {1-δ}{2}.$$ If the amplitude $a$ belongs to the Hörmander class $S^{m_1}_{ρ,δ}$ and $ϕ\in Φ^{2}$ satisfies the strong non-degeneracy condition, then we prove that the following Fourier integral operator $T_{ϕ,a}$ defined by \begin{align*} T_{ϕ,a}f(x)=\int_{\mathbb{R}^{n}}e^{iϕ(x,ξ)}a(x,ξ)\widehat{f}(ξ)dξ, \end{align*} is bounded from the local Hardy space $h^1(\mathbb{R}^n)$ to $L^1(\mathbb{R}^n)$. As a corollary, we can also obtain the corresponding $L^p(\mathbb{R}^n)$-boundedness when $1<p<2$. These theorems are rigorous improvements on the recent works of Staubach and his collaborators. When $0\leq ρ\leq 1,δ\leq \max\{ρ,1-ρ\}$, by using some similar techniques in this note, we can get the corresponding theorems which coincide with the known results.

math.CA↗

Fourier integral operators on Hardy spaces with Hormander class

In this note, we consider a Fourier integral operator defined by \begin{align*} T_{ϕ,a}f(x) = \int_{\mathbb{R}^{n}}e^{iϕ(x,ξ)}a(x,ξ)\widehat{f} ξ)dξ, \end{align*}here $a$ is the amplitude, and $ϕ$ is the phase. Let $0\leqρ\leq 1,n\geq 2$ or $0\leqρ<1,n=1$ and $$m_p=\frac{ρ-n}{p}+(n-1)\min\{\frac 12,ρ\}.$$ If $a$ belongs to the forbidden Hörmander class $S^{m_p}_{ρ,1}$ and $ϕ\in Φ^{2}$ satisfies the strong non-degeneracy condition, then for any $\frac {n}{n+1}<p\leq 1$, we can show that the Fourier integral operator $T_{ϕ,a}$ is bounded from the local Hardy space $h^p$ to $L^p$. Furthermore, if $a$ has compact support in variable $x$, then we can extend this result to $0<p\leq 1$. As $S^{m_p}_{ρ,δ}\subset S^{m_p}_{ρ,1}$ for any $0\leq δ\leq 1$, our result supplements and improves upon recent theorems proved by Staubach and his collaborators for $a\in S^{m}_{ρ,δ}$ when $δ$ is close to 1. As an important special case, when $n\geq 2$, we show that $T_{ϕ,a}$ is bounded from $H^1$ to $L^1$ if $a\in S^{(1-n)/2}_{1,1}$ which is a generalization of the well-known Seeger-Sogge-Stein theorem for $a\in S^{(1-n)/2}_{1,0}$. This result is false when $n=1$ and $a\in S^{0}_{1,1}$.

math.DG↗

Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation

Multi-aspect controllable text generation aims to control the generated texts in attributes from multiple aspects (e.g., "positive" from sentiment and "sport" from topic). For ease of obtaining training samples, existing works neglect attribute correlations formed by the intertwining of different attributes. Particularly, the stereotype formed by imbalanced attribute correlations significantly affects multi-aspect control. In this paper, we propose MAGIC, a new multi-aspect controllable text generation method with disentangled counterfactual augmentation. We alleviate the issue of imbalanced attribute correlations during training using counterfactual feature vectors in the attribute latent space by disentanglement. During inference, we enhance attribute correlations by target-guided counterfactual augmentation to further improve multi-aspect control. Experiments show that MAGIC outperforms state-of-the-art baselines in both imbalanced and balanced attribute correlation scenarios. Our source code and data are available at https://github.com/nju-websoft/MAGIC.

cs.CL↗

$L^p$ boundedness of pseudo-differential operators with symbols in $S^{n(ρ-1)/2}_{ρ,1}$

For symbol $a\in S^{n(ρ-1)/2}_{ρ,1}$ the pseudo-differential operator $T_a$ may not be $L^2$ bounded. However, under some mild extra assumptions on $a$, we show that $T_a$ is bounded from $L^{\infty}$ to $BMO$ and on $L^p$ for $2\leq p<\infty$. A key ingredient in our proof of the $L^{\infty}$-$BMO$ boundedness is that we decompose a cube, use $x$-regularity of the symbol and combine certain $L^2$, $L^\infty$ and $L^{\infty}$-$BMO$ boundedness. We use an almost orthogonality argument to prove an $L^2$ boundedness and then interpolation to obtain the desired $L^p$ boundedness.

math.CA↗

Heterogeneous Federated Knowledge Graph Embedding Learning and Unlearning

Federated Learning (FL) recently emerges as a paradigm to train a global machine learning model across distributed clients without sharing raw data. Knowledge Graph (KG) embedding represents KGs in a continuous vector space, serving as the backbone of many knowledge-driven applications. As a promising combination, federated KG embedding can fully take advantage of knowledge learned from different clients while preserving the privacy of local data. However, realistic problems such as data heterogeneity and knowledge forgetting still remain to be concerned. In this paper, we propose FedLU, a novel FL framework for heterogeneous KG embedding learning and unlearning. To cope with the drift between local optimization and global convergence caused by data heterogeneity, we propose mutual knowledge distillation to transfer local knowledge to global, and absorb global knowledge back. Moreover, we present an unlearning method based on cognitive neuroscience, which combines retroactive interference and passive decay to erase specific knowledge from local clients and propagate to the global model by reusing knowledge distillation. We construct new datasets for assessing realistic performance of the state-of-the-arts. Extensive experiments show that FedLU achieves superior results in both link prediction and knowledge forgetting.

cs.LG↗

Entropic destruction of heavy quarkonium in a rotating hot and dense medium from holography

Previous studies have indicated that the peak of the quarkonium entropy at the deconfinement transition can be related to the entropic force which would induce the dissociation of heavy quarkonium. In this paper, we study the entropic force in a rotating hot and dense medium using AdS/CFT correspondence. It turns out that the inclusion of angular velocity increases the entropic force thus enhancing quarkonium dissociation, while chemical potential has the same effect. Furthermore, the results imply that the quarkonium dissociates easier in rotating medium compared to static case.

nucl-th↗

Some notes on endpoint estimates for pseudo-differential operators

We study the pseudo-differential operator \begin{equation*} T_a f\left(x\right)=\int_{\mathbb{R}^n}e^{ix\cdotξ}a\left(x,ξ\right)\widehat{f}\left(ξ\right)\,\textrm{d}ξ, \end{equation*} where the symbol $a$ is in the Hörmander class $S^{m}_{ρ,1}$ or more generally in the rough Hörmander class $L^{\infty}S^{m}_ρ$ with $m\in\mathbb{R}$ and $ρ\in [0,1]$. It is known that $T_a$ is bounded on $L^1(\mathbb{R}^n)$ for $m<n(ρ-1)$. In this paper we mainly investigate its boundedness properties when $m$ is equal to the critical index $n(ρ-1)$. For any $0\leq ρ\leq 1$ we construct a symbol $a\in S^{n(ρ-1)}_{ρ,1}$ such that $T_a$ is unbounded on $L^1$ and furthermore it is not of weak type $(1,1)$ if $ρ=0$. On the other hand we prove that $T_a$ is bounded from $H^1$ to $L^1$ if $0\leq ρ<1$ and construct a symbol $a\in S^0_{1,1}$ such that $T_a$ is unbounded from $H^1$ to $L^1$. Finally, as a complement, for any $1<p<\infty$ we give an example $a\in S^{-1/p}_{0,1}$ such that $T_a$ is unbounded on $L^p(\mathbb{R})$.

math.CA↗