arXiv · 2503.17365
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
Abstract
Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models: DeepSeek-R1-8B, Gemma-2-9B, Llama 3.1-8B, and Qwen2.5-7B. We show that while Llama-based models exhibited significant harm reduction through self-critique, other architectures demonstrated less improvement in harm detection after abliteration. These results suggest CAI's effectiveness may vary depending on model architecture and reasoning capabilities.
Explore related subjects
Keep this discovery
Antonio-Gabriel Chacón Menke, Phan Xuan Tan. 2025-02-01. How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers. https://arxiv.org/abs/2503.17365
Cite the original work for its findings. Save a collection to share your selection of sources.