Searcharxiv⌕ Search

arXiv · 2609.22723

ANI-Gamut: Benchmarking Agent Reliability across the Gamut of Agent-Network Interface Abstractions

Abstract

Large Language Model (LLM) agents are increasingly trusted to operate live networks: they read state, change configuration, and verify the result. A first-order question is left implicit: at which level of abstraction should the agent operate? We make the interface-abstraction level an explicit, controlled experimental variable, organizing agent-network interfaces into a spectrum from raw CLI (A0) through bounded wrappers (A1) and standardized model-driven configuration (A2) to typed transactional service intent (A3-T), reconciled source-of-truth automation (A3-R), and their combination (A4). We present ANI-Gamut, a reproducible, open-source playground that exposes the same task at several levels on a single, densely populated brownfield substrate, where many coexisting services share resources and collateral damage actually arises. We instantiate and measure four points of the spectrum and describe the others only at the conceptual level, and record three dependent variables as the substrate is stressed by injected faults: task reliability, collateral damage against pre-existing tenants, and operational cost. In a pilot with small run counts (n=20, n=6 and n=4 per level), the interface level moves reliability and cost sharply: on a live change-set task a raw-shell agent fails on all six seeds while a typed transactional interface succeeds on all six (paired McNemar p=0.031, on six discordant pairs), at roughly an order of magnitude less cost, as the engineering effort migrates from the agent to a reusable transaction layer. Collateral damage, by contrast, is absent at every level, whether benign or under faults: in a tenant-isolated substrate the agents fail safe, and the blast radius is held by the substrate's isolation, which moves the safety question from the agent to the substrate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lorenzo Bracciale, Pierpaolo Loreti, Andrea Mayer, Stefano Salsano, Wim Henderickx. 2026-09-19. ANI-Gamut: Benchmarking Agent Reliability across the Gamut of Agent-Network Interface Abstractions. https://arxiv.org/abs/2609.22723

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

SHORTCUT: In-Collective Topology Reconfiguration for Low-Latency AllReduce

Distributed ML training relies on efficient AllReduce communication to aggregate data across nodes. In this setting, reconfigurable optical interconnects offer high-bandwidth, energy-efficient direct links between accelerators but often produce ring-based topologies that remain static during a collective. The Ring AllReduce algorithm naturally matches these topologies but its cumulative latency grows linearly with node count. Low-latency algorithms such as Recursive Doubling (RD) instead achieve logarithmic cumulative per-step latency, but their long-distance exchanges incur dilation and congestion costs on a static ring. In-collective topology reconfiguration can eliminate these penalties but each topology change adds reconfiguration delay. The key question to improve AllReduce completion time is therefore not only how to reconfigure RD efficiently, but when selectively reconfigured RD becomes faster than the topology-matched Ring algorithm. We present Shortcut: an effective strategy for topology reconfiguration that enables RD to shortcut costly multi-hop communication only when it pays off - beyond the performance of the Ring algorithm. Across small and medium messages on 32 nodes, Shortcut achieves $4.4\times$-$6.0\times$ speedups over Ring. At 128 nodes, it remains up to $7\times$ faster with a $10\,μs$ reconfiguration delay, showing that selective reconfiguration is especially effective in latency-sensitive, large-scale settings.

cs.NI↗

Enhancing BGP Security by Understanding BGP's Language with LLMs

The trust-based nature of Border Gateway Protocol (BGP) makes it vulnerable to prefix hijacking and misconfigurations. Traditional BGP anomaly detection relies on manual inspection with poor scalability, while Machine/Deep Learning (M/DL)-based approaches suffer from suboptimal precision, limited generalizability, and high retraining cost. This is because existing M/DL methods focus on topological structures rather than semantic characteristics of Autonomous Systems (ASes), assigning dissimilar embeddings to functionally similar but topologically distant ASes. To address this, we propose BGPShield, a novel anomaly detection framework built on an Adaptive LLM BGP Encoder that captures each AS's Behavior Portrait and Routing Policy Rationale beyond topology. Inspired by multimodal LLMs, the encoder generates embeddings representing both routing behaviors and semantics of ASes via contrastive learning. We further introduce SAM-ED to quantify BGP-specific semantic deviations between historical and updated paths, rather than naively accumulating distances without awareness of BGP-specific structures. Evaluated on 16 real-world datasets, BGPShield detects 100% of verified anomalies with an average false discovery rate below 5%. The open-source LLMs used by BGPShield were released prior to several evaluation events, verifying generalizability on unseen events. Furthermore, BGPShield can construct the representation for a previously unseen AS within one second, significantly outperforming BEAM which demands thorough retraining (averagely 65 hours).

cs.NI↗

GATE: GPU-Accelerated Traffic Engineering for the WAN

Traffic engineering (TE) has become a crucial tool for enforcing routing policy and maintaining operational efficiency in large networks. Existing TE solutions pick an objective function to optimize, aiming to balance (i) allocating traffic optimally with (ii) reacting quickly to demand changes and disruption events. However, as the scale of networks grows, the runtime of the existing optimal solution becomes infeasibly large. The alternative - approximate solvers - result in costly inefficiencies. We present GPU-Accelerated Traffic Engineering (GATE), which achieves the best of both worlds: enabling fast TE runtimes through a highly-parallelizable GPU-compatible decomposition, while iteratively converging to the provably optimal solution. GATE unlocks a unique set of desirable properties: it becomes increasingly parallelizable with network size, supports a wide spectrum of fairness objectives, and offers theoretically guaranteed convergence to the optimal solution and near-optimal convergence within a bounded time. We evaluate GATE on production traces from two large cloud WANs, and show that GATE achieves near-optimal solutions 4-10x faster than state-of-the-art.

cs.NI↗