arXiv · 2410.15651
Understanding and Alleviating Memory Consumption in RLHF for LLMs
Abstract
Fine-tuning with Reinforcement Learning with Human Feedback (RLHF) is essential for aligning large language models (LLMs). However, RLHF often encounters significant memory challenges. This study is the first to examine memory usage in the RLHF context, exploring various memory management strategies and unveiling the reasons behind excessive memory consumption. Additionally, we introduce a simple yet effective approach that substantially reduces the memory required for RLHF fine-tuning.
Explore related subjects
Keep this discovery
Jin Zhou, Hanmei Yang, Steven, Tang, Mingcan Xiang, Hui Guan, Tongping Liu. 2024-10-21. Understanding and Alleviating Memory Consumption in RLHF for LLMs. https://arxiv.org/abs/2410.15651
Cite the original work for its findings. Save a collection to share your selection of sources.