Searcharxiv⌕ Search

arXiv subjects

Shuangshuang Tian

Publications and source records attributed to Shuangshuang Tian.

3 recordsLinked to original sources

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning

Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) dominate the post-training landscape for mathematical reasoning, yet differ fundamentally in their reliance on expert trajectories. To understand the optimal way to harness these trajectories for maximizing performance, we propose the Plasticity-Ceiling Framework. This framework empirically grounds the post-training landscape by decomposing the final performance ceiling into the foundational SFT performance and the subsequent RL plasticity (i.e., the maximum improvement via RL). Through extensive benchmarking, we establish the Sequential SFT-then-RL pipeline as the superior standard, overcoming the stability and premature convergence deficits inherent in synchronized approaches. Furthermore, we derive precise scaling guidelines: (1) Transitioning to RL at the Stable or Mild Overfitting Regime of SFT maximizes the final ceiling by securing a robust SFT foundation with substantial RL plasticity; (2) Refuting the ``Less is More'' hypothesis in SFT-then-RL scaling, we demonstrate that Data Scale determines the primary post-training potential, while Trajectory Difficulty acts as a performance multiplier; and (3) The Minimum Validation Loss of SFT serves as a reliable indicator for selecting the expert trajectories that maximize the ultimate performance ceiling. Our findings provide actionable guidelines for extracting maximum value from expert trajectories.

cs.LG↗

GlobalRAG: Enhancing Global Reasoning in Multi-hop Question Answering via Reinforcement Learning

Reinforcement learning has recently shown promise in improving retrieval-augmented generation (RAG). Despite these advances, its effectiveness in multi-hop question answering (QA) remains limited by two fundamental limitations: (i) global planning absence to structure multi-step reasoning, and (ii) unfaithful execution, which hinders effective query formulation and consistent use of retrieved evidence. We propose GlobalRAG, a reinforcement learning framework designed to enhance global reasoning in multi-hop QA. GlobalRAG decomposes questions into subgoals, coordinates retrieval with reasoning, and refines evidence iteratively. To guide this process, we introduce Planning Quality Reward and SubGoal Completion Reward, which encourage coherent planning and reliable subgoal execution. In addition, a progressive weight annealing strategy balances process-oriented and outcome-based objectives. Extensive experiments on both in-domain and out-of-domain benchmarks demonstrate that GlobalRAG significantly outperforms strong baselines while using only 8k training data (42% of the training data used by strong baselines), achieving average improvements of 14.2% in both EM and F1.

cs.CL↗

Mechanism of O$_2$ influence on the decomposition process of the eco-friendly gas insulating medium C$_4$F$_7$N/CO$_2$

The C$_4$F$_7$N/CO$_2$/O$_2$ gas mixture is the most promising eco-friendly gas insulation medium available. However, there are few studies on the mechanism of the influence of the buffer gas O2 ratio and its role in the decomposition characteristics of C4F7N/CO2. In this paper, based on the ReaxFF reaction molecular dynamics method and density functional theory, a simulation of the thermal decomposition process of the C$_4$F$_7$N/CO$_2$ mixture under different O2 ratios was carried out at temperatures in the range 2000-3000 K. A constructed model of the C4F7N/CO2/O2 mixture reaction system was used that included the possible reaction paths, product distribution characteristics and their generation rates. The calculation results show that the thermal decomposition of C$_4$F$_7$N/CO$_2$/O$_2$ mainly generates species such as CF$_3$, CF$_2$, CF, F, C$_2$F$_5$, C$_2$F$_4$, C$_2$F$_2$, C$_3$F$_7$, C$_2$F$_2$N, C$_3$F$_4$N, CFN, CN, CO, O, and C. Among them, the two particles CF$_2$ and CN are the most abundant. The first decomposition time of C$_4$F$_7$N is advanced by the addition of O$_2$, while the amount of C$_4$F$_7$N decomposed and the generation of major decomposed particles decreases. The addition of 0%-4% of O$_2$ decreases the reaction rate of the main decomposition reaction in the reaction system. Quantum chemical calculations show that the dissociation process occurring from the combination of C$_4$F$_7$N with O atom is more likely to occur compared to the direct dissociation process of C$_4$F$_7$N molecules. The conclusions of this study provide a theoretical basis for the optimization of the application ratio of C$_4$F$_7$N/CO$_2$/O$_2$ and the diagnosis of its equipment operation and maintenance.

physics.comp-ph↗