SearcharxivSearch

arXiv subjects

Zi-Han Wang

Publications and source records attributed to Zi-Han Wang.

9 recordsLinked to original sources

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.

cs.AI

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related instances into progressively challenging sequences and uses a single policy to alternate between task solving and curating an evolving skill document passed directly to the next task. Decoupled credit assignment across tasks supervises solving with the current task outcome and curation with discounted downstream outcomes. Experiments on ALFWorld, WebShop, and ScienceWorld show that SkillRise achieves the strongest Pass@1 performance among the compared methods, with gains over the strongest baseline ranging from 2.3 to 8.5 percentage points. Although trained across distinct tasks, its learned curation policy remains effective for repeated attempts on the same task. Further analysis reveals scaling at test time across tasks: performance improves with longer sequences of related tasks even when each task is attempted only once. This trend suggests that SkillRise reuses transferable skills across tasks rather than benefiting from repeated sampling of the same task. SkillRise further retains strong performance while substantially reducing the runtime overhead of skill learning pipelines with multiple stages. Together, these results provide a simple and efficient training paradigm for LLM agents to extract, refine, and reuse transferable skills across tasks.

cs.LG

Self-Distilled Agentic Reinforcement Learning

Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coarse supervision for long-horizon interaction. On-Policy Self-Distillation (OPSD) complements RL by introducing dense token-level guidance from a teacher branch augmented with privileged context. However, transferring OPSD to multi-turn agents proves problematic: compounding multi-turn instability destabilizes supervision, while skill-conditioned privileged guidance requires asymmetric treatment for negative teacher rejections may arise from imperfect skills retrieval or utilization. We introduce SDAR (Self-Distilled Agentic Reinforcement Learning), which treats OPSD as a gated auxiliary objective while keeping RL as the primary optimization backbone. SDAR maps detached token-level signals into a sigmoid gate, strengthening distillation on teacher-endorsed positive-gap tokens and softly attenuating negative teacher rejections. Across the Qwen2.5 and Qwen3 families on ALFWorld, WebShop, and Search-QA, SDAR substantially improves over GRPO (+9.4% on ALFWorld, +7.0% on Search-QA, +10.2% on WebShop-Acc), avoids the instability of naive GRPO+OPSD, and consistently outperforms hybrid RL--OPSD baselines across model scales.

cs.LG

The Klein bottle ratio of two-dimensional ferromagnetic Potts models

The weakly first-order nature of the two-dimensional 5-state ferromagnetic Potts model poses challenges for numerical study. Using density-matrix and tensor-network renormalization group methods, we investigate these transitions of the Potts-$q$ model via the Klein bottle ratio $g$ on original and dual lattices. Finite-size scaling of $g$ as a function of transverse system size $L_y$ accurately locates the critical points for $q = 4, 5, 6$. We further examine the transfer-matrix spectra and entanglement entropy, extracting central charges through toroidal and Klein bottle boundary conditions. For $q = 5$, the extracted central charge ($c \approx 1.14811$) is close to the real part of the theoretical value $c_{5\text{-Potts}} = 1.1375 \pm 0.0211 i$ predicted by complex conformal field theories. The observed drift in the scaling exponent $b$ effectively distinguishes the continuous transition from the weakly first-order regime. Furthermore, the extrapolated divergence of $g$ confirms the first-order nature of the $q=5$ Potts model.

cond-mat.stat-mech

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges

The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve. However, the progress of such systems is hindered by the lack of rigorous, quantitative evaluation. To tackle this challenge, we introduce CreativeBench, a benchmark for evaluating machine creativity in code generation, grounded in a classical cognitive framework. Comprising two subsets -- CreativeBench-Combo and CreativeBench-Explore -- the benchmark targets combinatorial and exploratory creativity through an automated pipeline utilizing reverse engineering and self-play. By leveraging executable code, CreativeBench objectively distinguishes creativity from hallucination via a unified metric defined as the product of quality and novelty. Our analysis of state-of-the-art models reveals distinct behaviors: (1) scaling significantly improves combinatorial creativity but yields diminishing returns for exploration; (2) larger models exhibit ``convergence-by-scaling,'' becoming more correct but less divergent; and (3) reasoning capabilities primarily benefit constrained exploration rather than combination. Finally, we propose EvoRePE, a plug-and-play inference-time steering strategy that internalizes evolutionary search patterns to consistently enhance machine creativity.

cs.AI

Systematic investigation on the superheavy nucleus formation in the reactions of $^{48}$Ca, $^{50}$Ti, $^{51}$V and $^{54}$Cr on actinide nuclei

The synthesis of superheavy elements strongly relies on the competition of the quasifission and fusion fission dynamics in the fusion-evaporation reactions. The systematics on the formation of superheavy nuclei in the $^{48}$Ca, $^{50}$Ti, $^{51}$V and $^{54}$Cr induced fusion reactions on actinide nuclei $^{232}$Th, $^{231}$Pa, $^{238}$U, $^{237}$Np, $^{242,244}$Pu, $^{243}$Am, $^{245,248}$Cm, $^{249}$Bk, $^{249}$Cf has been thoroughly investigated with the dinuclear system model by including the cluster transfer and coupling to the dynamical evolution of the quadrupole deformation parameters. The uncertainties of the fusion-evaporation excitation functions with the mass models of FRDM2012, KTUY05, LDM1966, SkyHFB, WS4 are investigated and compared with the available experimental data from Dubna, GSI, Berkeley and RIKEN. The production cross sections, optimal evaporation channels and beam energies in the synthesis of superheavy elements Z = 119 and 120 were predicted and compared for the different mass models in the reactions of $^{50}\mathrm{Ti} + ^{249}\mathrm{Bk}$, $^{51}\mathrm{V} + ^{248}\mathrm{Cm}$, $^{54}\mathrm{Cr} + ^{243}\mathrm{Am}$, $^{50}\mathrm{Ti} + ^{249}\mathrm{Cf}$, $^{51}\mathrm{V} + ^{249}\mathrm{Bk}$, $^{54}\mathrm{Cr} + ^{248}\mathrm{Cm}$, respectively.

nucl-th

Dynamics of light nuclei produced in the massive transfer reactions

Within the framework of the dinuclear system (DNS) model by implementing the cluster transfer into the dissipation process, we systematically investigated the energy spectra and the angular distribution of the preequilibrium clusters (n, p, d, t, $^{3}$He, $\alpha$, $^{6,7}$Li, $^{8,9}$Be) in the massive transfer reactions of $^{12}$C+$^{209}$Bi, $^{14}$N+$^{159}$Tb, $^{14}$N+$^{169}$Tm, $^{14}$N+$^{181}$Ta, $^{14}$N+$^{197}$Au, $^{14}$N+$^{209}$Bi, $^{58,64,72}$Ni+$^{198}$Pt near the Coulomb barrier energies. It is found that the neutron emission is the most probable in comparison with the charged particles and the $\alpha$ yields are comparable with the hydrogen isotopes in magnitude. The preequilibrium clusters are mainly produced from the projectile-like and target-like fragments in the evolution of dinuclear system. The kinetic energy spectra manifest the Boltzmann distribution and the Coulomb potential influences the structure. The preequilibrium clusters follows the angular distribution of multinucleon transfer fragments.

nucl-th

Power-law distribution and scale-invariant structure from the first CHIME/FRB Fast Radio Burst catalog

We study the statistical property of fast radio bursts (FRBs) based on a selected sample of 190 one-off FRBs in the first CHIME/FRB catalog. Three power law models are used in the analysis, and we find the cumulative distribution functions of energy can be well fitted by bent power law and thresholded power law models. And the distribution functions of fluctuations of energy well follow the Tsallis $q$-Gaussian distribution. The $q$ values in the Tsallis $q$-Gaussian distribution are constant with small fluctuations for different temporal scale intervals, indicating a scale-invariant structure of the bursts. The earthquakes and soft gamma repeaters show similar properties, which are consistent with the predictions of self-organized criticality systems.

astro-ph.HE

Production cross-sections of new superheavy elements with Z = 119-120 in fusion-evaporation reactions

We have calculated production cross sections of new superheavy elements with atomic number Z=119,120 in the fusion-evaporation reactions of $^{48}$Ca+$^{252}$Es, $^{48}$Ca+$^{257}$Fm, $^{49}$Sc+$^{252}$Es, $^{49}$Sc+$^{251}$Cf, $^{50}$Ti+$^{247}$Bk, $^{50}$Ti+$^{251}$Cf, $^{51}$V+$^{247}$Cm, $^{51}$V+$^{247}$Cf, $^{54}$Cr+$^{243}$Am, $^{54}$Cr+$^{247}$Cm, $^{56}$Mn+$^{244}$Pu, $^{56}$Mn+$^{243}$Am, $^{60}$Fe+$^{237}$Np, $^{60}$Fe+$^{244}$Pu, $^{61}$Co+$^{238}$U, $^{61}$Co+$^{237}$Np, $^{64}$Ni+$^{231}$Pa, $^{64}$Ni+$^{238}$U, $^{65}$Cu+$^{232}$Th, $^{65}$Cu+$^{231}$Pa, and $^{68}$Zn+$^{232}$Th within the dinuclear system model systematically. The inner fusion barriers have been extracted from the driving potential or potential energy surface which could be used to predict the relative fusion probability roughly. The influence of mass asymmetry of the colliding partners on the production of new superheavy elements (SHE) has been investigated systematically. It is found that fusion probability increase along with the increasing mass asymmetry of colliding systems. The Ti-induced reactions have the largest cross-sections of the new SHE. The dependence of production cross-sections of new superheavy elements on the isospin of projectile nuclei has been discussed. The new SHE of $^{289-295}$119 has been predicted as the synthesis cross sections around serval picobarns in the $^{46,48,50,52}$Ti-induced reactions. Production cross-section of the element of $^{295}$120 has been evaluated as large as 1 picobarn in the reactions $^{46}$Ti ($^{251}$Cf, 2n) $^{295}$120 at $E^*$ = 26 MeV. The optimal projectile-target combinations and beam energies for producing new SHE with atomic number Z=119-120 are proposed for the forthcoming experiments.

nucl-th