arXiv · 2607.02975
Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve
Abstract
Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, achieved the highest social-access success at 0.65, a 0.11 decrease from centralized indirect access with an interval that included zero. Exploratory comparisons separated five of six model pairs under social access, while neither centralized setting separated any at this sample size. In post-hoc task-blocked tests, three pairwise interaction $p$-values remained significant after multiplicity adjustment. A reference-relative reader associates the wider gaps with failures to obtain needed information. Because the data cannot distinguish ineffective requests by the evaluated agent from inaccurate replies by simulated participants, the reader's labels describe where trajectories stopped, not why models differed.
Explore related subjects
Keep this discovery
Dan C. Hsu, Luke Lu. 2026-08-30. Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve. https://arxiv.org/abs/2607.02975
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.