arXiv · 2509.24418
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
Abstract
As large language models (LLMs) are integrated into numerous applications, LLMs' safety becomes critical for both application developers and intended users. Currently, great efforts have been made to develop safety benchmarks with fine-grained taxonomies. However, these benchmarks' taxonomies are disparate with different safety policies. Thus, existing safeguards trained on these benchmarks are either coarse-grained to only distinguish between "safe'' and "unsafe,'' or constrained by the specified narrow risk taxonomies. To leverage these fine-grained safety policies across multiple safety taxonomies, we propose GSPR, a Generalizable Safety Policy Reasoner to identify unsafe inputs and outputs with violated safety taxonomies and concise explanations. Unlike prior safeguards which only cover a fixed set of risk factors, GSPR incentivizes its reasoning capability with varied safety taxonomies through reinforcement learning. Our GSPR can be trained across multiple safety benchmarks with distinct taxonomies and naturally exhibits powerful generalization ability. We conduct extensive experiments to show that GSPR significantly improves existing safety guardrails' reasoning capabilities for both safety and category prediction tasks. Moreover, GSPR also achieves the least inference token costs with explanations.
Explore related subjects
Keep this discovery
Haoran Li, Jingru Zeng, Yulin Chen, Huihao Jing, Wenbin Hu, Hao Peng, Haochen Shi, Xi Yang, Ziqian Zeng, Sirui Han, Yangqiu Song. 2026-08-28. GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners. https://arxiv.org/abs/2509.24418
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.