TY - RPRT TI - Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions AU - Federico Bianchi AU - Mirac Suzgun AU - Giuseppe Attanasio AU - Paul Röttger AU - Dan Jurafsky AU - Tatsunori Hashimoto AU - James Zou PY - 2024 UR - https://arxiv.org/abs/2309.07875 ID - 2309.07875 ER -