arXiv · 2601.06366
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
Abstract
Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction.
Explore related subjects
Keep this discovery
Pratyush Desai, Luoxi Tang, Yuqiao Meng, Zhaohan Xi. 2026-01-10. SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use. https://arxiv.org/abs/2601.06366
Cite the original work for its findings. Save a collection to share your selection of sources.