arXiv · 2406.11880
Knowledge Return Oriented Prompting (KROP)
Abstract
Many Large Language Models (LLMs) and LLM-powered apps deployed today use some form of prompt filter or alignment to protect their integrity. However, these measures aren't foolproof. This paper introduces KROP, a prompt injection technique capable of obfuscating prompt injection attacks, rendering them virtually undetectable to most of these security measures.
Explore related subjects
Keep this discovery
Jason Martin, Kenneth Yeung. 2024-06-11. Knowledge Return Oriented Prompting (KROP). https://arxiv.org/abs/2406.11880
Cite the original work for its findings. Save a collection to share your selection of sources.