TY - RPRT TI - Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards AU - Alexandre Ramé AU - Guillaume Couairon AU - Mustafa Shukor AU - Corentin Dancette AU - Jean-Baptiste Gaya AU - Laure Soulier AU - Matthieu Cord PY - 2023 UR - https://arxiv.org/abs/2306.04488 ID - 2306.04488 ER -