What a Policy Gate Can and Cannot Know: Measured Boundaries of Cross-Platform Command Adjudication
Gateways that adjudicate an agent's actions before they execute are only as good as their understanding of the action. We study a policy gate that never parses shell syntax: it consumes a typed, realised action (verb, operands, resolved zones, program-object identity) and decides ALLOW, ASK or DENY. Working on a Linux twin of the Windows benchmark of our previous study [1], we ask how faithfully it adjudicates, what survives translation, and whether deciding stays affordable as the system is used. A frozen 61-case table scores 61/61 in two rounds with no false allow; a 50-operator mutation campaign kills 46 of 50 mutants (92.0%), with all four surviving mutants classified. An 82-row audit yields 43 re-expressions, 21 carrier differences, 16 study-specific inapplicable rows and two unresolved cases; of 25 rows labelled "no counterpart," four remain unmapped here. Consulted by the executor, 25 escapes become zero with no benign payload blocked. In an exploratory one-gateway snapshot, three unauthenticated endpoint labels produce case-level non-refusal majorities of 81.8-98.0%, but immediate-execution majorities of 4.0-52.5%. Adjudication reads no accumulating state; credential verification does, scanning its whole ledger. We report that cost and the fix we would make.