arXiv · 2609.34862
JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges
Abstract
Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative judges on four benchmark collections totaling 5,219 trajectories, using a common risk rubric and behavior-level labels. JEV attains a benchmark-averaged positive-class F1 of 77.8, compared with 74.1 for the strongest generative configuration, GLM-5.2, with valid-result coverage of 95.5\% and 94.4\%, respectively. Performance varies across datasets, with JEV leading on ATBench500 and MCPHunt and GLM leading on R-Judge and TraceSafe. Across the four benchmarks, JEV's median successful-call latency is 0.99 seconds; estimated token cost averages \$0.000195 per valid judgment. These results support JEV as an economical screening signal, with trade-offs in precision and recall.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhiqiang Wang, Yichao Gao. 2026-09-28. JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges. https://arxiv.org/abs/2609.34862
Cite the original work for its findings. Save a collection to share your selection of sources.