TY - RPRT TI - Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems AU - Yicheng Cai AU - Mitchell John DeStefano AU - Guodong Dong AU - Pulkit Handa AU - Peng Liu AU - Tejas Singhal AU - Peiyu Tseng AU - Winston Jen White PY - 2026 UR - https://arxiv.org/abs/2603.28998 ID - 2603.28998 ER -