SearcharxivSearch

arXiv subjects

Yizhu Liu

Publications and source records attributed to Yizhu Liu.

10 recordsLinked to original sources

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to distill these traces into reusable feedback without post-hoc outcome labels, drawing on their evidence of local progress, recovery, and unfinished requirements. We introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes this evidence into evidence-grounded nested shortcut trees. DENSE compresses redundant attempts, reconciles issues across levels using recovery evidence, and summarizes completed branches while expanding unresolved ones, linking reusable progress to remaining obligations. We introduce REFIT, a source-paired protocol comparing feedback from shared initial trajectories under post-hoc outcome blindness, with environments and model contexts reset for fresh attempts at the same tasks. On Terminal-Bench 2.1, DENSE achieves the highest strict pass rate among tested non-privileged feedback methods across four recipient models. Relative to initial executions, strict pass rate improves by 7.12-15.64 pp, with 19.0-43.6% fewer observed recipient tokens in reruns. GPT-5.5 ablations support combining nested subtask analysis with shortcut construction and issue reconciliation. These findings point toward agent self-refinement through evidence-grounded trajectory reuse with less reliance on external supervision.

cs.AI

Improving Topic Relevance Model by Mix-structured Summarization and LLM-based Data Augmentation

Topic relevance between query and document is a very important part of social search, which can evaluate the degree of matching between document and user's requirement. In most social search scenarios such as Dianping, modeling search relevance always faces two challenges. One is that many documents in social search are very long and have much redundant information. The other is that the training data for search relevance model is difficult to get, especially for multi-classification relevance model. To tackle above two problems, we first take query concatenated with the query-based summary and the document summary without query as the input of topic relevance model, which can help model learn the relevance degree between query and the core topic of document. Then, we utilize the language understanding and generation abilities of large language model (LLM) to rewrite and generate query from queries and documents in existing training data, which can construct new query-document pairs as training data. Extensive offline experiments and online A/B tests show that the proposed approaches effectively improve the performance of relevance modeling.

cs.IR

ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art

Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs). In this work, we select ASCII art as a representative artifact. It depicts concepts through careful arrangement of characters, which can be formulated in both text and image modalities. We frame the problem as a recognition task, and construct a novel benchmark, ASCIIEval. It covers over 3K samples with an elaborate categorization tree, along with a training set for further enhancement. Encompassing a comprehensive analysis of tens of models through different input modalities, our benchmark demonstrate its multi-faceted diagnostic power. Given textual input, language models shows their visual perception ability on ASCII art concepts. Proprietary models achieve over 70% accuracy on certain categories, with GPT-5 topping the rank. For image inputs, we reveal that open-source MLLMs suffer from a trade-off between fine-grained text recognition and collective visual perception. They exhibit limited generalization ability to this special kind of arts, leading to the dramatic gap of over 20.01% accuracy compared with their proprietary counterparts. Another critical finding is that model performance is sensitive to the length of the ASCII art, with this sensitivity varying across input modalities. Unfortunately, none of the models could successfully benefit from the simultaneous provision of both modalities, highlighting the need for more flexible modality-fusion approaches. Besides, we also introduce approaches for further enhancement and discuss future directions. Resources are available at https://github.com/JiaQiSJTU/VisionInText.

cs.CL

Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model

Despite tremendous improvements in natural language generation, summarization models still suffer from the unfaithfulness issue. Previous work evaluates faithfulness either using models trained on the other tasks or in-domain synthetic data, or prompting a large model such as ChatGPT. This paper proposes to do zero-shot faithfulness evaluation simply with a moderately-sized foundation language model. We introduce a new metric FFLM, which is a combination of probability changes based on the intuition that prefixing a piece of text that is consistent with the output will increase the probability of predicting the output. Experiments show that FFLM performs competitively with or even outperforms ChatGPT on both inconsistency detection and faithfulness rating with 24x fewer parameters. FFLM also achieves improvements over other strong baselines.

cs.CL

Multi-turn Response Selection using Dialogue Dependency Relations

Multi-turn response selection is a task designed for developing dialogue agents. The performance on this task has a remarkable improvement with pre-trained language models. However, these models simply concatenate the turns in dialogue history as the input and largely ignore the dependencies between the turns. In this paper, we propose a dialogue extraction algorithm to transform a dialogue history into threads based on their dependency relations. Each thread can be regarded as a self-contained sub-dialogue. We also propose Thread-Encoder model to encode threads and candidates into compact representations by pre-trained Transformers and finally get the matching score through an attention layer. The experiments show that dependency relations are helpful for dialogue context understanding, and our model outperforms the state-of-the-art baselines on both DSTC7 and DSTC8*, with competitive results on UbuntuV2.

cs.CL

Taxonomy of Abstractive Dialogue Summarization: Scenarios, Approaches and Future Directions

Abstractive dialogue summarization is to generate a concise and fluent summary covering the salient information in a dialogue among two or more interlocutors. It has attracted great attention in recent years based on the massive emergence of social communication platforms and an urgent requirement for efficient dialogue information understanding and digestion. Different from news or articles in traditional document summarization, dialogues bring unique characteristics and additional challenges, including different language styles and formats, scattered information, flexible discourse structures and unclear topic boundaries. This survey provides a comprehensive investigation on existing work for abstractive dialogue summarization from scenarios, approaches to evaluations. It categorizes the task into two broad categories according to the type of input dialogues, i.e., open-domain and task-oriented, and presents a taxonomy of existing techniques in three directions, namely, injecting dialogue features, designing auxiliary training tasks and using additional data.A list of datasets under different scenarios and widely-accepted evaluation metrics are summarized for completeness. After that, the trends of scenarios and techniques are summarized, together with deep insights on correlations between extensively exploited features and different scenarios. Based on these analyses, we recommend future directions including more controlled and complicated scenarios, technical innovations and comparisons, publicly available datasets in special domains, etc.

cs.CL

In-sample Curriculum Learning by Sequence Completion for Natural Language Generation

Curriculum learning has shown promising improvements in multiple domains by training machine learning models from easy samples to hard ones. Previous works which either design rules or train models for scoring the difficulty highly rely on task-specific expertise, and cannot generalize. Inspired by the "easy-to-hard" intuition, we propose to do in-sample curriculum learning for natural language generation tasks. Our learning strategy starts training the model to generate the last few words, i.e., do sequence completion, and gradually extends to generate the whole output sequence. Comprehensive experiments show that it generalizes well to different tasks and achieves significant improvements over strong baselines.

cs.CL

Post-Training Dialogue Summarization using Pseudo-Paraphrasing

Previous dialogue summarization techniques adapt large language models pretrained on the narrative text by injecting dialogue-specific features into the models. These features either require additional knowledge to recognize or make the resulting models harder to tune. To bridge the format gap between dialogues and narrative summaries in dialogue summarization tasks, we propose to post-train pretrained language models (PLMs) to rephrase from dialogue to narratives. After that, the model is fine-tuned for dialogue summarization as usual. Comprehensive experiments show that our approach significantly improves vanilla PLMs on dialogue summarization and outperforms other SOTA models by the summary quality and implementation costs.

cs.CL

Origin of vibrational wavepacket dynamics in Fe carbene photosensitizer determined with femtosecond X-ray emission and scattering

Disentangling the dynamics of electrons and nuclei during nonadiabatic molecular transformations remains a considerable experimental challenge. Here we have investigated photoinduced electron transfer dynamics following a metal-to-ligand charge-transfer (MLCT) excitation of the [Fe(bmip)2]2+ photosensitizer, where bmip = 2,6-bis(3-methyl-imidazole-1- ylidine)-pyridine, with simultaneous femtosecond-resolution Fe Kα and K\b{eta} X-ray Emission Spectroscopy (XES) and Wide Angle X-ray Scattering (WAXS). This measurement clearly shows temporal oscillations in the XES and WAXS difference signals with the same 278 fs period oscillation. The oscillatory signal originates from an Fe-ligand stretching mode vibrational wavepacket on a triplet metal-centered (3MC) excited state surface. The vibrational wavepacket is created by 40% of the excited population that undergoes electron transfer from the non-equilibrium MLCT excited state to the 3MC excited state with a 110 fs time constant, while the other 60% relaxes to a 3MLCT excited state in parallel. The sensitivity of the Kα XES spectrum to molecular structure results from core-level vibronic coupling, due to a 0.7% average Fe-ligand bond length difference in the lowest energy geometry of the 1s and 2p core-ionized states. These results highlight the importance of vibronic effects in time-resolved XES experiments and demonstrate the role of metal-centered excited states in the electronic excited state relaxation dynamics of an Fe carbene photosensitizer.

physics.chem-ph

Femtosecond X-Ray Scattering Study of Ultrafast Photoinduced Structural Dynamics in Solvated [Co(terpy)2]2+

We study the structural dynamics of photoexcited [Co(terpy)2]2+ in an aqueous solution with ultrafast x-ray diffuse scattering experiments conducted at the Linac Coherent Light Source. Through direct comparisons with density functional theory calculations, our analysis shows that the photoexcitation event leads to elongation of the Co-N bonds, followed by coherent Co-N bond length oscillations arising from the impulsive excitation of a vibrational mode dominated by the symmetrical stretch of all six Co-N bonds. This mode has a period of 0.33 ps and decays on a subpicosecond time scale. We find that the equilibrium bond-elongated structure of the high spin state is established on a single-picosecond time scale and that this state has a lifetime of ~ 7 ps.

physics.chem-ph