arXiv · 2608.01454
How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection
Abstract
Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols. We investigate this dependency by re-evaluating representative PIDS on public datasets that meet our audit, labeling, and calibration requirements. Focusing primarily on the audited DARPA TC E3 datasets, we apply a unified protocol with temporally separated test periods and validation-only checkpoint selection and threshold calibration, and ask which architectural claims are empirically supported. We find that alerting success and investigation utility can diverge sharply, as several systems surface attacks without providing enough process-level context to support forensic investigation. Across the four primary datasets, a simple allowlist built from executable names and paths observed during training matches or exceeds the selected learned baselines on key operating-point metrics, showing that comparable performance on these metrics is achievable using lexical novelty alone. Quantifying semantic signal quality through feature completeness and field entropy helps explain why several audited E3 datasets support alerting performance without reliably separating model architectures. In contrast, Theia provides the richest semantic signal and shows the clearest improvements in ranking and node-level recovery for our reference model. Overall, these findings reinforce the importance of interpreting architectural claims in PIDS together with the benchmark properties and evaluation protocol that produced them.
Explore related subjects
Keep this discovery
Lorenzo Guerra, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi, Van-Tam Nguyen. 2026-08-02. How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection. https://arxiv.org/abs/2608.01454
Cite the original work for its findings. Save a collection to share your selection of sources.