arXiv · 2609.33532
Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection
Abstract
Reproducing a provenance-based intrusion detector's score does not establish what that score says about its emitted alarms or the information its encoder uses. We audit nine released implementations, execute four detectors using their own code, and isolate three measurement effects. First, a fixed-alert comparison separates label choice from neighbourhood credit: ThreaTrace reports precision 0.938 with neighbourhood credit, although only seven of its 994 alarms carry its own attack label. Second, removing test-label checkpoint selection lowers attack detection precision (ADP) by 0.16 to 0.33 across four forty-member word2vec configurations without changing detector order. Third, a buffer-reuse defect gives a linear encoder unintended degree-dependent inputs. Correcting it lowers type-only ADP in every identical-input initialization pair on two hosts, while historical word2vec effects depend on the host. These findings qualify claims from the inspected implementations about alarm precision, performance magnitude and static-attribute sufficiency. STRICT connects them to six checkable reporting requirements. Because the comparisons condition on benchmark targets, they neither validate those labels nor establish a universal detector ranking.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jihwan Moon, Gunhee Kim, Myeongjang Pyeon. 2026-09-27. Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection. https://arxiv.org/abs/2609.33532
Cite the original work for its findings. Save a collection to share your selection of sources.