SearcharxivSearch

arXiv subjects

Zhiyu Cao

Publications and source records attributed to Zhiyu Cao.

At least 19 recordsLinked to original sources

Hint-Guided Diversified Policy Optimization for LLM Reasoning

Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising enhancement strategy. However, existing reward mechanisms are constrained to the outcome-level correctness and lack explicit signals to guide the model to consider diverse solutions. In contrast, human problem solving typically involves evaluating multiple potential approaches and selecting the most reliable solution, a cognitive process that current RLVR frameworks do not explicitly incentivize. Inspired by this, we propose Hint-Guided Diversified Policy Optimization (HDPO), allowing the model to first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning. HDPO comprises two stages of Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning to incentivize the model to generate diverse and reliable solutions following the ``propose-select-think'' trajectory. Experimental results show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the LLM's ability to identify reliable solutions.

cs.CL

Discourse Coherence and Response-Guided Context Rewriting for Multi-Party Dialogue Generation

Previous research on multi-party dialogue generation has predominantly leveraged structural information inherent in dialogues to directly inform the generation process. However, the prevalence of colloquial expressions and incomplete utterances in dialogues often impedes comprehension and weakens the fidelity of dialogue structure representations, which is particularly pronounced in multi-party dialogues. In this work, we propose a novel framework DRCR (Discourse coherence and Response-guided Context Rewriting) to improve multi-party dialogue generation through dialogue context rewriting. Specifically, DRCR employs two complementary feedback signals, discourse coherence and response quality, to construct preference data for both context rewriting and response generation. Moreover, we propose a dynamic self-evolution learning method that allows the rewriter and responder to continuously enhance their capabilities through mutual interaction in an iterative training loop. Comprehensive experiments conducted on four multi-party dialogue datasets substantiate the effectiveness of DRCR.

cs.CL

Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search

Conversational Query Rewriting (CQR) aims to rewrite ambiguous queries to achieve more efficient conversational search. Early studies have predominantly focused on the rewriting in isolation, ignoring the feedback from query rewrite, passage retrieval and response generation in the rewriting process. To address this issue, we propose Multi-Faceted Self-Consistent Preference Aligned CQR (MSPA-CQR). Specifically, we first construct self-consistent preference alignment data from three dimensions (rewriting, retrieval, and response) to generate more diverse rewritten queries. Then we propose prefix guided multi-faceted direct preference optimization to learn preference information from three different dimensions. The experimental results show that our MSPA-CQR is effective in both in- and out-of-distribution scenarios.

cs.CL

Statistics of Thermal Avalanches in Driven Amorphous Systems

Within the framework of the random first-order transition theory of glasses, we discuss the statistics of thermal avalanches, the large scale rearrangements in driven amorphous systems near their instability. Stringy excitations yield nonPoisson waiting time statistics. Embedding these statistics in a generalized Master equation captures the nonMarkovian, aging dynamics of avalanche clusters. We apply this framework to analyze nonequilibrium signatures of thermal avalanches, auto correlation functions and effective temperatures, under both quasi static shear and stochastic shaking protocols. We use full counting statistics to derive the complete distribution of both the avalanche magnitudes and avalanche counts, uncovering the intermediate time behavior.

cond-mat.dis-nn

Steering Active-Colloid Assembly by Biasing Dissipation

Complex nonequilibrium self-assembly enables the formation of materials with specific patterns and functions from the bottom up. How to directionally control the assembly to form the target configuration is a challenge. Here, we propose a dissipation bias principle for targeted assembly, which highlights that controlling the dissipation tendency can play an important role by modulating the frequency and intensity of local rearrangements. Following this principle, one can induce ordered target configurations from disordered structures and also achieve directional selection among multiple assembly pathways. We use the assembly of active colloids as a platform to show our results.

cond-mat.soft

Universal trade-off between irreversibility and intrinsic timescale in thermal relaxation with applications to thermodynamic inference

We establish a general lower bound for the entropy production rate (EPR) based on the Kullback-Leibler divergence and the Logarithmic-Sobolev constant that characterizes the time-scale of relaxation. This bound can be considered as an enhanced second law of thermodynamics. When applied to thermal relaxation, it reveals a universal trade-off relation between the dissipation rate and the intrinsic relaxation timescale. From this relation, a thermodynamic upper bound on the relaxation time between two given states emerges, acting as an inverse speed limit over the entire time region. We also obtain a quantum version of this upper bound, which is always tighter than its classical counterpart, incorporating an additional term due to decoherence. Remarkably, we further demonstrate that the trade-off relation remains valid for any generally non-Markovian coarse-grained relaxation dynamics, highlighting its significant applications in thermodynamic inference. This trade-off relation is a new tool in inferring EPRs in molecular dynamics simulations and practical experiments.

cond-mat.stat-mech

ICR: Iterative Clarification and Rewriting for Conversational Search

Most previous work on Conversational Query Rewriting employs an end-to-end rewriting paradigm. However, this approach is hindered by the issue of multiple fuzzy expressions within the query, which complicates the simultaneous identification and rewriting of multiple positions. To address this issue, we propose a novel framework ICR (Iterative Clarification and Rewriting), an iterative rewriting scheme that pivots on clarification questions. Within this framework, the model alternates between generating clarification questions and rewritten queries. The experimental results show that our ICR can continuously improve retrieval performance in the clarification-rewriting iterative process, thereby achieving state-of-the-art performance on two popular datasets.

cs.CL

Incomplete Utterance Rewriting with Editing Operation Guidance and Utterance Augmentation

Although existing fashionable generation methods on Incomplete Utterance Rewriting (IUR) can generate coherent utterances, they often result in the inclusion of irrelevant and redundant tokens in rewritten utterances due to their inability to focus on critical tokens in dialogue context. Furthermore, the limited size of the training datasets also contributes to the insufficient training of the IUR model. To address the first issue, we propose a multi-task learning framework EO-IUR (Editing Operation-guided Incomplete Utterance Rewriting) that introduces the editing operation labels generated by sequence labeling module to guide generation model to focus on critical tokens. Furthermore, we introduce a token-level heterogeneous graph to represent dialogues. To address the second issue, we propose a two-dimensional utterance augmentation strategy, namely editing operation-based incomplete utterance augmentation and LLM-based historical utterance augmentation. The experimental results on three datasets demonstrate that our EO-IUR outperforms previous state-of-the-art (SOTA) baselines in both open-domain and task-oriented dialogue. The code will be available at https://github.com/Dewset/EO-IUR.

cs.CL

Two-stage Incomplete Utterance Rewriting on Editing Operation

Previous work on Incomplete Utterance Rewriting (IUR) has primarily focused on generating rewritten utterances based solely on dialogue context, ignoring the widespread phenomenon of coreference and ellipsis in dialogues. To address this issue, we propose a novel framework called TEO (\emph{Two-stage approach on Editing Operation}) for IUR, in which the first stage generates editing operations and the second stage rewrites incomplete utterances utilizing the generated editing operations and the dialogue context. Furthermore, an adversarial perturbation strategy is proposed to mitigate cascading errors and exposure bias caused by the inconsistency between training and inference in the second stage. Experimental results on three IUR datasets show that our TEO outperforms the SOTA models significantly.

cs.CL

Motorized Chromosome Models of Mitosis

During mitosis, near-spherical chromosomes reconfigure into rod-like structures to ensure their accurate segregation to daughter cells. We explore here, the interplay between the nonequilibrium activity of molecular motors in determining the chromosomal organization in mitosis and its characteristic symmetry-breaking events. We present a hybrid motorized chromosome model that highlights the distinct roles of condensin I and II in shaping mitotic chromosomes. Guided by experimental observations, the simulations suggest that condensin II facilitates large-scale scaffold formation, while condensin I is paramount in local helical loop arrangement. Together, these two distinct grappling motors establish the hierarchical helical structure characteristic of mitotic chromosomes, which exhibit striking local and, sometimes global, chirality and contribute to the robust mechanical properties of mitotic chromosomes. Accompanying the emergence of rigidity, the model provides mechanisms of forming defects, including perversions and entanglements, and shows how these may be partially resolved through condensin activity and topoisomerase action. This framework bridges coarse-grained energy landscape models of chromosome dynamics and non-equilibrium molecular dynamics, advancing the understanding of chromosome organization during cell division and beyond.

physics.bio-ph

Price-mediated contagion with endogenous market liquidity

Price-mediated contagion occurs when a positive feedback loop develops following a drop in asset prices which forces banks and other financial institutions to sell their holdings. Prior studies of such events fix the level of market liquidity without regards to the level of stress applied to the system. This paper introduces a framework to understand price-mediated contagion in a system where the capacity of the market to absorb liquidated assets is determined endogenously. In doing so, we construct a joint clearing system in interbank payments, asset prices, and market liquidity. We establish mild assumptions which guarantee the existence of greatest and least clearing solutions. We conclude with detailed numerical case studies which demonstrate the, potentially severe, repercussions of endogenizing the market liquidity on system risk.

q-fin.RM

Large Language Model in Financial Regulatory Interpretation

This study explores the innovative use of Large Language Models (LLMs) as analytical tools for interpreting complex financial regulations. The primary objective is to design effective prompts that guide LLMs in distilling verbose and intricate regulatory texts, such as the Basel III capital requirement regulations, into a concise mathematical framework that can be subsequently translated into actionable code. This novel approach aims to streamline the implementation of regulatory mandates within the financial reporting and risk management systems of global banking institutions. A case study was conducted to assess the performance of various LLMs, demonstrating that GPT-4 outperforms other models in processing and collecting necessary information, as well as executing mathematical calculations. The case study utilized numerical simulations with asset holdings -- including fixed income, equities, currency pairs, and commodities -- to demonstrate how LLMs can effectively implement the Basel III capital adequacy requirements. Keywords: Large Language Models, Prompt Engineering, LLMs in Finance, Basel III, Minimum Capital Requirements, LLM Ethics

q-fin.RM

Modeling Inverse Demand Function with Explainable Dual Neural Networks

Financial contagion has been widely recognized as a fundamental risk to the financial system. Particularly potent is price-mediated contagion, wherein forced liquidations by firms depress asset prices and propagate financial stress, enabling crises to proliferate across a broad spectrum of seemingly unrelated entities. Price impacts are currently modeled via exogenous inverse demand functions. However, in real-world scenarios, only the initial shocks and the final equilibrium asset prices are typically observable, leaving actual asset liquidations largely obscured. This missing data presents significant limitations to calibrating the existing models. To address these challenges, we introduce a novel dual neural network structure that operates in two sequential stages: the first neural network maps initial shocks to predicted asset liquidations, and the second network utilizes these liquidations to derive resultant equilibrium prices. This data-driven approach can capture both linear and non-linear forms without pre-specifying an analytical structure; furthermore, it functions effectively even in the absence of observable liquidation data. Experiments with simulated datasets demonstrate that our model can accurately predict equilibrium asset prices based solely on initial shocks, while revealing a strong alignment between predicted and true liquidations. Our explainable framework contributes to the understanding and modeling of price-mediated contagion and provides valuable insights for financial authorities to construct effective stress tests and regulatory policies.

q-fin.CP

Dynamic and Thermodynamic Origins of Motility-Induced Phase Separation

Active matter systems are inherently out of equilibrium and break the detailed balance (DB) at the microscopic scale, exhibiting vital collective phenomena such as motility-induced phase separation (MIPS). Here, we introduce a coarse-grained mapping method to probe DB breaking in the density-energy phase space, which allows us to reveal the dynamic and thermodynamic origins of MIPS based on nonequilibrium potential and flux landscape theory. Hallmarks of nonequilibrium properties are manifested by identifying the visible probability flux in the coarse-grained phase space. Remarkably, the flux for the system with the activity lower than the MIPS threshold tends to ``tear up" the single potential well of the uniform-density phase to create two wells of phases with different densities, presenting directly that the nonequilibrium flux is the dynamic origin of MIPS. Moreover, we find that the obtained entropy production rate (EPR) of the system undergoes a transition from nearly independent of activity to increasing proportionally as activity increases after the single well is "teared up". The transition of EPR's scaling behavior might provide a hint of the thermodynamic origin of MIPS in the coarse-grained space. Our findings propose a new route to explore the nonequilibrium nature of active systems, and provide new insights into dynamic and thermodynamic properties of MIPS.

cond-mat.soft

Fast Functionalization with High Performance in the Autonomous Information Engine

Mandal and Jarzynski have proposed a fully autonomous information heat engine, consisting of a demon, a mass and a memory register interacting with a thermal reservoir. This device converts thermal energy into mechanical work by writing information to a memory register, or conversely, erasing information by consuming mechanical work. Here, we derive a speed limit inequality between the relaxation time of state transformation and the distance between the initial and final distributions, where the combination of the dynamical activity and entropy production plays an important role. Such inequality provides a hint that a speed-performance trade-off relation exists between the relaxation time to functional state and the average production. To obtain fast functionalization while maintaining the performance, we show that the relaxation dynamics of information heat engine can be accelerated significantly by devising an optimal initial state of the demon. Our design principle is inspired by the so-called Mpemba effect, where water freezes faster when initially heated.

cond-mat.stat-mech

Designing Autonomous Maxwell Demon via Stochastic Resetting

Autonomous Maxwell demon is a new type of information engine proposed by Mandal and Jarzynski, which can produce work by exploiting an information tape. Here, we show that a stochastic resetting mechanism can be used to improve the performance of autonomous Maxwell demons notably. Generally, the performance is composed of two important features, the time cost for an autonomous demon to reach its functional state and its efficacious working region in its functional state. Here, we provide a set of design principles for the system, which are capable of improving the two important features. On the one hand, one can drive any autonomous demon system to its functional periodic steady state at a fastest pace for any initial distribution through resetting the demon for a predetermined critical time and closing the reset after that. On the other hand, the system can reach a new functional state when the resetting is always on, in which case the efficacious region of the demon being extended significantly. Moreover, a dual function region in a new phase diagram of the demon with resetting has been found. Remarkably, in this dual function region the demon with resetting can realize anomalous output of work and erasure of information on the tape simultaneously, violating the second law of thermodynamics apparently. To this question, we derive a new modified Clausius inequality to restore the second law by taking the cost of resetting into account.

cond-mat.stat-mech

Improved estimation for energy dissipation in biochemical oscillations

Biochemical oscillations, regulating the timing of life processes, need consume energy to achieve good performance on crucial functions, such as high accuracy of phase period and high sensitivity to external signals. However, it is a great challenge to precisely estimate the energy dissipation in such systems. Here, based on the stochastic normal form theory (SNFT), we calculate the Pearson correlation coefficient between the oscillatory amplitude and phase, and a trade-off relation between transport efficiency and phase sensitivity can then be derived, which serves as a tighter form than the estimator resulting from the conventional thermodynamic uncertainty relation (TUR). Our findings demonstrate that a more precise energy dissipation estimation can be obtained by enhancing the sensitivity of the biochemical oscillations. Moreover, the internal noise and amplitude power effects have also been discovered.

cond-mat.stat-mech

Effective Entropy Production and Thermodynamic Uncertainty Relation of Active Brownian Particles

Understanding stochastic thermodynamics of active Brownian particles (ABPs) system has been an important topic in very recent years. In this article we study a general model of active Brownian particle systems by introducing a coarse-grained Fokker-Planck equation, which allows us to identify an effective entropy production along a stochastic trajectory, wherein an activity and configuration dependent diffusion coefficient comes into play with an important role. Although the hidden component between the true entropy production and the effective one is dominant, the effective entropy production still act as a reliable measure to quantify the dynamical irreversibility, capturing important phenomenon such as the interface and defects of motility induced phase separation (MIPS). Furthermore, in this framework, we are able to obtain the entropic bound as well as TUR associated with any generalized currents in the systems. We expect the new conceptual quantities proposed here to be broadly used in the context of active matter.

cond-mat.stat-mech