SearcharxivSearch

arXiv subjects

Xunlong Wang

Publications and source records attributed to Xunlong Wang.

4 recordsLinked to original sources

Self-Improving Large Language Models via Progressive Experience Evolution

Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training-time optimization methods can update model parameters but lack an explicit mechanism for accumulating transferable experience. Bridging these two paradigms requires a critical intermediate stage that remains underexplored, namely \emph{experience distillation}. To address this gap, we propose \textbf{SPEE} (\textbf{S}elf-\textbf{P}rogressive \textbf{E}xperience \textbf{E}volution), a unified post-training framework that sequentially performs explicit experience evolution followed by implicit policy optimization. During explicit experience evolution, SPEE reflects on trajectories collected from multiple interactions to extract, verify, and progressively evolve transferable experience, which is subsequently internalized into the policy through privilege-guided On-Policy Self-Distillation (OPSD). During implicit policy optimization, reward-driven reinforcement learning leverages these internalized priors to explore novel solution strategies. In the experience evolution stage, a continuously evolving global experience pool consolidates knowledge from both successful and failed trajectories, filters out low-utility experience, and mitigates post-hoc rationalization induced by individual trajectories. Experiments on five mathematical reasoning benchmarks demonstrate that SPEE consistently outperforms both test-time and training-time self-evolution baselines across three model scales. The source code is available at https://github.com/rrrsj/SPEE.

cs.CL

Nonlinear contagion dynamics on dynamical networks: exact solutions ranging from consensus times to evolutionary trajectories

Understanding nonlinear social contagion dynamics on dynamical networks, such as opinion formation, is crucial for gaining new insights into consensus and polarization. Similar to threshold-dependent complex contagions, the nonlinearity in adoption rates poses challenges for mean-field approximations. To address this theoretical gap, we focus on nonlinear binary-opinion dynamics on dynamical networks and analytically derive local configurations, specifically the distribution of opinions within any given focal individual's neighborhood. This exact local configuration of opinions, combined with network degree distributions, allows us to obtain exact solutions for consensus times and evolutionary trajectories. Our counterintuitive results reveal that neither biased assimilation (i.e., nonlinear adoption rates) nor preferences in local network rewiring -- such as in-group bias (preferring like-minded individuals) and the Matthew effect (preferring social hubs) -- can significantly slow down consensus. Among these three social factors, we find that biased assimilation is the most influential in accelerating consensus. Furthermore, our analytical method efficiently and precisely predicts the evolutionary trajectories of adoption curves arising from nonlinear contagion dynamics. Our work paves the way for enabling analytical predictions for general nonlinear contagion dynamics beyond opinion formation.

physics.soc-ph

Opinion dynamics on biased dynamical networks: beyond rare opinion updating

Opinion dynamics is of paramount importance as it provides insights into the complex dynamics of opinion propagation and social relationship adjustment. It is assumed in most of the previous works that social relationships evolve much faster than opinions. This is not always true in reality. We propose an analytical approximation to study this issue for arbitrary time scales between opinion adjustment and network evolution. To this end, the coefficient of determination in statistics is introduced and a one-dimensional stable manifold is analytically found, i.e., the most likely trajectory. With the aid of the stable manifold, we further obtain the fate of opinions and the consensus time, i.e., fixation probability and fixation time. We find that for in-group bias, the more likely individuals are to adopt the popular opinion, the less likely the majority opinion takes over the population, i.e., conformity inhibits the domination of popular opinions. This counter-intuitive result can be interpreted from a game perspective, in which in-group bias refers to a coordination game and rewiring probability refers to a rescaling of the selection intensity. Our work proposes an efficient approximation method to foster the understanding of opinion dynamics in dynamical networks.

physics.soc-ph

A robust way to speed up consensus via adaptive social networks

Opinion dynamics is crucial for unraveling the complexities of human interaction in the information age. How to speed up consensus without disturbing the fate of the system is key for opinion dynamics. We propose a voter model on adaptive networks, which resembles the coevolutionary process between opinions and social relationships. We prove the existence of a one-dimensional stable manifold for the system, which facilitates us to study both the fate of the system and the consensus time it takes. Surprisingly, we find the adjustment of social relations speeds up consensus but does not affect the fate of the system. For echo-chamber-like networks which consist of two homogeneous subnetworks connected by few sparse links, a small probability of adaptive edge dynamics is sufficient to accelerate consensus formation, which is counterintuitive. If the network structure makes consensus much slower than that of the regular networks, minor random rewiring makes a discontinuous drop in consensus time. Our work opens up an avenue for speeding up consensus without disturbing the fate of the system. It can be insightful for crowd control.

physics.soc-ph