SearcharxivSearch

arXiv subjects

Bradley Monton

Publications and source records attributed to Bradley Monton.

3 recordsLinked to original sources

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document constrains its behavior over an extended tool-use horizon. We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks. Each task places an agent in a self-contained company environment (a file workspace with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol) and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20-124 pages. Tasks span five domains (finance, medical billing, insurance, logistics, and HR) and 10 fictional companies. To resist memorization, every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies. Grading is fully deterministic: each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not. Under strict grading, where a trial passes only if every criterion is satisfied, the strongest evaluated model passes 36.2% of trials, and most frontier models remain below 25%. Failures follow consistent patterns: agents let a plausible but unauthorized in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve. We release the tasks, environments, and evaluation harness.

cs.AI

Counting Marbles with 'Accessible' Mass Density: A Reply to Bassi and Ghirardi

In a previous article (cf. quant-ph/9905065) we argued that, while Lewis is correct that the enumeration principle fails in dynamical wavepacket reduction theories, one need not following Lewis in rejecting these theories. Because the dynamical reduction process itself prevents the failure of enumeration from ever becoming manifest, and because one can treat the semantics for dynamical reduction theories as not adding anything of ontological import to them, it is reasonable to accept these theories notwithstanding Lewis's counting anomaly. In their response to our paper (cf. quant-ph/9907050), Bassi and Ghirardi reject our criticisms of their own response to Lewis, as well as our argument against Lewis that dynamical reduction precludes failures of enumeration from ever becoming manifest. Our intention here is to demonstrate that Bassi and Ghirardi's responses to us do not succeed.

quant-ph

Losing Your Marbles in Wavefunction Collapse Theories

Peter Lewis ([1997]) has recently argued that the wavefunction collapse theory of GRW (Ghirardi, Rimini, and Weber [1986]) can only solve the problem of wavefunction tails at the expense of predicting that arithmetic does not apply to ordinary macroscopic objects. More specifically, Lewis argues that the GRW theory must violate the enumeration principle: that `if marble 1 is in the box and marble 2 is in the box and so on through marble $n$, then all $n$ marbles are in the box' ([1997], p. 321). Ghirardi and Bassi ([1999]) have replied that it is meaningless to say that the enumeration principle is violated because the wavefunction Lewis uses to exhibit the violation cannot persist, according to the GRW theory, for more than a split second ([1999], p. 709). On the contrary, we argue that Lewis's argument survives Ghirardi and Bassi's criticism unscathed. We then go on to show that, while the enumeration principle can fail in the GRW theory, the theory itself guarantees that the principle can never be empirically falsified, leaving the applicability of arithmetical reasoning to both micro- and macroscopic objects intact.

quant-ph