arXiv · 2610.00759
Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents
Abstract
Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: https://anonymous.4open.science/r/RL-Transfer-between-env-4F47/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sabrina Saika, Yinuo Du, Aritran Piplai. 2026-09-30. Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents. https://arxiv.org/abs/2610.00759
Cite the original work for its findings. Save a collection to share your selection of sources.