arXiv · 2609.12839
Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
Abstract
The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing an escalating risk as these models can bypass proprietary API guardrails when deployed locally. However, SLMs deployed as autonomous agents often struggle with long-horizon, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges due to context bloat and cognitive degradation from accumulated tool-call outputs. To understand and mitigate this cybersecurity threat, we introduce context segmentation, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems. Evaluating on the picoCTF dataset using memory-constrained gemma-4 models, we demonstrate that for the E4B model, our strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries, and successfully solving 18.52% of tasks that standard agentic execution fails to complete. Code is available at https://github.com/9xeb/context-segmentation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sebastiano Nordio, Michele Lotto. 2026-09-15. Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks. https://arxiv.org/abs/2609.12839
Cite the original work for its findings. Save a collection to share your selection of sources.