arXiv · 2601.22475
Continual Policy Consolidation for Lifelong Robot Learning
Abstract
Building a generalist robot policy requires continuously integrating new skills while preserving previously acquired behaviors. Directly optimizing a single policy over a growing task stream is difficult because robotic interaction is expensive, task distributions are heterogeneous, and sequential updates induce interference. To address these problems, we propose continual policy consolidation (CPC), a teacher--student framework that combines continual policy distillation with prioritized experience replay and expandable experts. This architecture separates skill acquisition from policy consolidation: teachers are trained independently through reinforcement learning, and their behaviors are continually distilled into a central generalist student. This decomposition retains the practical strength of reinforcement learning for task-specialized training while casting student-side consolidation as a supervised policy-learning problem. To balance stability and plasticity as the task stream grows, the student combines an expandable Transformer-based mixture-of-experts architecture with prioritized trajectory replay. Extensive experiments show that the student recovers a large proportion of teacher performance while achieving near-zero forgetting. These results demonstrate a scalable route for consolidating independently acquired robot skills into a continually growing generalist policy.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qijun He, Yuxuan Li, Mingqi Yuan, Xiaoquan Sun, Wen-Tse Chen, Jeff Schneider, Jiayu Chen. 2026-01-30. Continual Policy Consolidation for Lifelong Robot Learning. https://arxiv.org/abs/2601.22475
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.