TY - RPRT TI - How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective AU - Teng Xiao AU - Mingxiao Li AU - Yige Yuan AU - Huaisheng Zhu AU - Chao Cui AU - Vasant G Honavar PY - 2024 UR - https://arxiv.org/abs/2410.10093 ID - 2410.10093 ER -