arXiv · 2602.12566
To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) plays a key role in stimulating the explicit reasoning capability of Large Language Models (LLMs). We can achieve expert-level performance in some specific domains via RLVR, such as coding or math. When a general multi-domain expert-level model is required, we need to carefully consider the collaboration of RLVR across different domains. The current state-of-the-art models mainly employ two different training paradigms for multi-domain RLVR: mixed multi-task RLVR and separate RLVR followed by model merging. However, most of the works did not provide a detailed comparison and analysis about these paradigms. To this end, we choose multiple commonly used high-level tasks (e.g., math, coding, science, instruction following, and agent) as our target domains and design extensive qualitative and quantitative experiments using open-source datasets. We find the RLVR across domains exhibits small mutual interferences, and reasoning-intensive domains have mutually synergistic effects. Furthermore, we analyze the internal mechanisms from the perspectives of information constraints, model prediction behavior and self-verification. Our homepage is at https://github.com/Mosi-AI/M2RL.
Explore related subjects
Keep this discovery
Haoqing Wang, Xiang Long, Ziheng Li, Yilong Xu, Tingguang Li, Yehui Tang. 2026-02-13. To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models. https://arxiv.org/abs/2602.12566
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.