SearcharxivSearch

arXiv subjects

Zongye Zhang

Publications and source records attributed to Zongye Zhang.

11 recordsLinked to original sources

LivingRAG: Augmenting Graph RAG with Experience

Graph-based RAG improves multi-hop question answering by organizing evidence as a knowledge graph. However, most existing RAG systems process each query in isolation and discard useful reasoning from the LLM's response after inference. As a result, later related queries need to retrieve evidence and reason from scratch. We propose LivingRAG, a Graph RAG framework with writable and reusable reasoning experience. LivingRAG adds a writable experience store to a graph-based retrieval backbone, enabling verified experiences to be reused during inference in two ways. Stored graph signals help retrieval find entities and passages that were useful in earlier related queries. Stored summaries provide a reference reasoning pattern for answer generation. We analyze online QA streams and find reusable signals from shared entities, graph neighborhoods, and question templates. Experiments on multi-hop QA benchmarks show that LivingRAG improves accuracy over strong RAG baselines and reduces completion-token use when relevant prior experience is reused.

cs.AI

Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation

Generalizing motion representation across diverse characters remains challenging due to significant topological variations in skeletal structures across datasets and species, which hinder the development of scalable generative models. To bridge this gap, we propose a Semantic-Aware Topology-Agnostic framework that learns a unified latent manifold shared by disparate species. Unlike methods relying on fixed hierarchies or rigid padding strategies, our approach leverages a semantic modulation mechanism to align functional joint correspondences, thereby decoupling motion from topology. This design enables the construction of a continuous, generative-friendly motion space from large-scale, unaligned raw BVH data. Experiments on human and animal datasets demonstrate that our framework achieves high-fidelity reconstruction and supports downstream text-to-motion tasks. Notably, the model enables zero-shot cross-species retargeting without paired data. Code and demos are available at: https://github.com/zzysteve/SATA

cs.GR

Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion

Generating 3D human motion from text descriptions remains challenging due to the diverse and complex nature of human motion. While existing methods excel within the training distribution, they often struggle with out-of-distribution motions, limiting their applicability in real-world scenarios. Existing VQVAE-based methods often fail to represent novel motions faithfully using discrete tokens, which hampers their ability to generalize beyond seen data. Meanwhile, diffusion-based methods operating on continuous representations often lack fine-grained control over individual frames. To address these challenges, we propose a robust motion generation framework MoMADiff, which combines masked modeling with diffusion processes to generate motion using frame-level continuous representations. Our model supports flexible user-provided keyframe specification, enabling precise control over both spatial and temporal aspects of motion synthesis. MoMADiff demonstrates strong generalization capability on novel text-to-motion datasets with sparse keyframes as motion prompts. Extensive experiments on two held-out datasets and two standard benchmarks show that our method consistently outperforms state-of-the-art models in motion quality, instruction fidelity, and keyframe adherence. The code is available at: https://github.com/zzysteve/MoMADiff

cs.CV

SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation

While current skeleton action recognition models demonstrate impressive performance on large-scale datasets, their adaptation to new application scenarios remains challenging. These challenges are particularly pronounced when facing new action categories, diverse performers, and varied skeleton layouts, leading to significant performance degeneration. Additionally, the high cost and difficulty of collecting skeleton data make large-scale data collection impractical. This paper studies one-shot and limited-scale learning settings to enable efficient adaptation with minimal data. Existing approaches often overlook the rich mutual information between labeled samples, resulting in sub-optimal performance in low-data scenarios. To boost the utility of labeled data, we identify the variability among performers and the commonality within each action as two key attributes. We present SkeletonX, a lightweight training pipeline that integrates seamlessly with existing GCN-based skeleton action recognizers, promoting effective training under limited labeled data. First, we propose a tailored sample pair construction strategy on two key attributes to form and aggregate sample pairs. Next, we develop a concise and effective feature aggregation module to process these pairs. Extensive experiments are conducted on NTU RGB+D, NTU RGB+D 120, and PKU-MMD with various GCN backbones, demonstrating that the pipeline effectively improves performance when trained from scratch with limited data. Moreover, it surpasses previous state-of-the-art methods in the one-shot setting, with only 1/10 of the parameters and much fewer FLOPs. The code and data are available at: https://github.com/zzysteve/SkeletonX

cs.CV

A Survey on Data Synthesis and Augmentation for Large Language Models

The success of Large Language Models (LLMs) is inherently linked to the availability of vast, diverse, and high-quality data for training and evaluation. However, the growth rate of high-quality data is significantly outpaced by the expansion of training datasets, leading to a looming data exhaustion crisis. This underscores the urgent need to enhance data efficiency and explore new data sources. In this context, synthetic data has emerged as a promising solution. Currently, data generation primarily consists of two major approaches: data augmentation and synthesis. This paper comprehensively reviews and summarizes data generation techniques throughout the lifecycle of LLMs, including data preparation, pre-training, fine-tuning, instruction-tuning, preference alignment, and applications. Furthermore, We discuss the current constraints faced by these methods and investigate potential pathways for future development and research. Our aspiration is to equip researchers with a clear understanding of these methodologies, enabling them to swiftly identify appropriate data generation strategies in the construction of LLMs, while providing valuable insights for future exploration.

cs.CL

Selected strong decays of pentaquark State $P_c(4312)$ in a chiral constituent quark model

The newly confirmed pentaquark state $P_c(4312)$ has been treated as a weakly bound $(Σ_c\bar{D})$ state by a well-established chiral constituent quark model and by a dynamical calculation on quark degrees of freedom where the quark exchange effect is accounted for. The obtained mass $4308$ MeV agrees with data. In this work, the selected strong decays of the $P_c(4312)$ state are studied with the obtained wave function. It is shown that the width of the $Λ_c\bar{D}^*$ decay is overwhelmed and the branching ratios of the $p\,η_c$ and $p\,J/ψ$ decays are both less than 1 percentage.

hep-ph

On the form factors of $d^*(2380)$

In order to explore the possible physical quantities for judging different structures of the newly observed resonance $d^*(2380)$, we study its electromagnetic form factors. In addition to the electric charge monopole $C0$, we calculate its electric quadrupole $E2$, magnetic dipole $M1$, and six-pole $M3$ form factors on the base of the realistic coupled $ΔΔ+CC$ channel $d^*$ wave function with both the $S$- and $D$-partial waves. The results show that the magnetic dipole moment and electric quadrupole deformation of $d^*$ are 7.602 and $2.53\times 10^{-2}~\rm{fm}^2$, respectively. The calculated magnetic dipole moment in the naive constituent quark model is also compared with the result of $D_{12}π$ picture. By comparing with partial results where the $d^*$ state is considered with a single $ΔΔ$ and with a $D_{12}π$ structures, we find that in addition to the charge distribution of $d^*(2380)$, the magnetic dipole moment and magnetic radius can be used to discriminate different structures of $d^*$. Moreover, a quite small electric quadrupole deformation indicates that $d^*$ is more inclined to an slightly oblate shape due to our compact hexaquark dominated structure of $d^*(2380)$.

hep-ph

On the charge distribution of $d^*(2380)$

We calculate the charge distributions of $d^*(2380)$. Two different interpretations of the $d^*$ are considered for a comparison. One is a compact explanation with coupled $ΔΔ+CC$ two-channel approximation in the chiral constituent quark model. Another is a resonance state of $D_{12}π$. The remarkable differences of the charge distributions in the two pictures are shown and it is expected that the future experiments may provide a clear test for the different theoretical interpretations.

nucl-th

Decay width of $d^*(2380) \to NN π$ process in a chiral constituent quark model

The width of three-body single-pion decay process $d^*\to NNπ^{0,\pm}$ is calculated by using the $d^*$ wave function obtained from our chiral SU(3) constituent quark model calculation. The effect of the dynamical structure on the width of $d^*$ is taken into account in both the single $ΔΔ$ channel and coupled $ΔΔ+CC$ two-channel approximations. Our numerical result shows that in the coupled-channel approximation, namely, the hidden-color configuration being considered, the obtained partial decay width of $d^*\to NNπ$ is about several hundred $\rm {KeV}$, while in the single $ΔΔ$ channel it is just about $2\sim 3~\rm{MeV}$. We, therefore, conclude that the partial width in the single-pion decay process of $d^*$ is much smaller than the widths in its double-pion decay processes. Our prediction may provide a criterion for judging different interpretations of the $d^*$ structure, as different pictures for the $d^*$ may result quite different partial decay width.

nucl-th

Decay width of $d^*(2380)\to NN ππ$ processes

The decay widths of four-body double-pion decays $\ds\to pn π^0π^0$, $\ds\to pn π^+π^-$, and iso-scalar parts of $\ds\to pp π^0π^-$ and $\ds\to nn π^+π^0$ are explicitly calculated with the help of the $d^*$ wave function obtained in a chiral SU(3) quark model calculation. The effect of the dynamical structure on $\ds$'s width is analyzed both in the single $ΔΔ$ channel and coupled $ΔΔ$ and $CC$ channel approximations. It is found that in the coupled-channel approximation, the obtained partial decay widths of $\ds\to pn π^0π^0$, $\ds\to pn π^+π^-$, and those of $d^*$ to the iso-scalar parts of $pp π^0π^-$ and $nn π^+π^0$ are about $7.4$MeV, $16.4$MeV, $3.5$MeV and $3.5$MeV, respectively As a consequence, the total width is about $64.5$MeV. These widths are consistent with those estimated by using the corresponding cross section data in our previous investigation and also the observed data. But in the single $ΔΔ$ channel approximation, the widths are still almost 2-times larger than the measured values. Apparently, the explicitly calculated width together with the evaluated mass of $d^*$ in the coupled $ΔΔ$ and $CC$ channel approximation can well explain the observed data, which again supports our assertion that the $\ds$ resonance is a six-quark dominated exotic state.

hep-ph

A study of $d^*(2380)\to d ππ$ decay width

The decay widths of the $\ds\to d π^0π^0$ and $\ds\to d π^+π^-$ processes are explicitly calculated in terms of our chiral quark model. By using the experimental ratios of cross sections between various decay channels, the partial widths of the $\ds\to pn π^0π^0$, $\ds\to pn π^+π^-$, $\ds\to pp π^0π^-$, and $\ds\to nn π^+π^0$ channels are also extracted. Further including the estimated partial width for the $\ds\to pn $ process, the total width of the $\ds$ resonance is obtained. In the first step of the practical calculation, the effect of the dynamical structure on the width of $\ds$ is studied in the single $ΔΔ$ channel approximation. It is found that the width is reduced by few tens of MeV, in comparison with the one obtained by considering the effect of the kinematics only. This presents the importance of such effect from the dynamical structure. However, the obtained width with the single $ΔΔ$ channel wave function is still too large to explain the data. It implies that the $\ds$ resonance will not consist of the $ΔΔ$ structure only, and instead there should be enough room for other structure such as the hidden-color (CC) component. Thus, in the second step, the width of $\ds$ is further evaluated by using a wave function obtained in the coupled $ΔΔ$ and CC channel calculation in the framework of the Resonating Group Method (RGM). It is shown that the resultant total width for $\ds$ is about 69 MeV, which is compatible with the experimental observation of about 75 MeV and justifies our assertion that the $\ds$ resonance is a hexaquark-dominated exotic state.

nucl-th