Searcharxiv⌕ Search

arXiv · 2610.05981

Mechanizing the User's Eye: Pre-Registered Deployment of a Sabotage-Validated Fail-Plausible Observer in a Production LLM Agent Runtime

Abstract

A prior longitudinal study of silent failures in a production LLM agent runtime (arXiv:2606.14589) found that about 70% were discovered by a human looking at the product as a user while thousands of tests and governance checks stayed green, and posed mechanizing part of what the human eye does as an open problem. This paper reports our attempt. We built an automated user-viewpoint observer targeting the most dangerous class, fail-plausible failure, in which an internal error becomes fluent, plausible output to the user. It is a two-layer pipeline: five deterministic signals distilled from incident postmortems escalate to an LLM judge whose verdicts must cite verbatim evidence or be discarded. Ground truth is 24 labeled production postmortems with explicit honesty boundaries: 16 of 24 are structurally invisible to any content-reading observer, and we say so. Offline, the deterministic layer achieves 6/6 regression detection with 0/4 false positives, each detector proven load-bearing by sabotage; held-out recall on novel patterns is 0/4. It is, so far, a regression engine. Deployment followed pre-registration: shadow mode caught and retired one systematic false positive on its first run, the 26-day shadow window then ran clean, flip criteria were fixed before the shadow data was read, and the analysis protocol was frozen before the enforcing-mode window opened. That window (12 observed days) fired zero verdicts; per the pre-registered path we report live precision as undefined rather than narrating quiet as success. The observer also produced ten silent failures of its own, confirming that the judge inherits the taxonomy it judges. We release the corpus, detector, and scorecard as a runnable bench: mechanization today retires the human's regression scanning so the eye can specialize in novelty; prediction remains open but is now measurable under frozen rules.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei Wu. 2026-10-05. Mechanizing the User's Eye: Pre-Registered Deployment of a Sabotage-Validated Fail-Plausible Observer in a Production LLM Agent Runtime. https://arxiv.org/abs/2610.05981

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Theory building in software engineering: Operationalization

This work is part of a research project whose ultimate goal is to systematize a procedure for constructing theories in software engineering. The proposed methodology involves four phases: conceptualization, operationalization, testing, and application. In previous work, we described the conceptualization process. This paper presents a set of procedures for systematizing the operationalization phase in theory building. Specifically, it translates the concepts and propositions obtained from the previous conceptualization into constructs and empirically testable hypotheses. The operationalization phase is structured across three distinct dimensions: the practical procedural steps, their strict formal mathematical specification, and their evaluation through an illustrative empirical case study. Grounded in critical realism, we operationalize the qualitatively derived concepts and propositions into explicit constructs and logical hypotheses. It is important to emphasize that a novel strategy has been designed that reduces the number of hypotheses resulting from this process. In the use case described, more than 30% of the initial hypotheses have been eliminated because they are redundant. This paper is a pioneering contribution in offering comprehensive guidelines for theory operationalization using logical implication and describes a strategy based on graph theory that tackles the problem of hypothesis explosion and maintains the principle of parsimony.

cs.SE↗

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes and returns a solver-feasible solution while encoding a semantically incorrect formulation. On compositional problems, the resulting feasibility-correctness gap reaches 90 percentage points. We introduce ReLoop, which combines two mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify) to reduce formulation errors during generation. Behavioral verification detects the errors that remain by testing whether the formulation responds correctly to solver-based parameter perturbation, a signal that comes from the solver rather than from LLM self-review and requires no ground truth. The two mechanisms address different error structures: structured generation gives the largest gain on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), and behavioral verification gives its largest gain on localized defects (+4.4pp on MAMO-ComplexLP). With diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6, and relative to direct generation it raises or preserves every reported metric of the three chat-tuned foundation models on all three benchmarks. For the narrowly fine-tuned SFT model we test, the chain-of-thought prompt conflicts with its learned output format and lowers its accuracy on MAMO-ComplexLP; we document and analyze this interaction. We release RetailOpt-190, 190 compositional retail optimization scenarios in which several constraints interact.

cs.SE↗

Finding Memory Leaks in C/C++ Programs via Neuro-Symbolic Augmented Static Analysis

Memory leaks remain prevalent in real-world C/C++ software. Static analyzers such as CodeQL provide scalable program analysis but frequently miss such bugs because they cannot recognize project-specific custom memory-management functions and lack path-sensitive control-flow modeling. We present MemHint, a neuro-symbolic pipeline that addresses both limitations by combining LLMs' semantic understanding of code with Z3-based symbolic reasoning. MemHint parses the target codebase and applies an LLM to classify each function as a memory allocator, deallocator, or neither, producing function summaries that record which argument or return value carries memory ownership, extending the analyzer's built-in knowledge beyond standard primitives such as malloc and free. A Z3-based validation step checks each summary against the function's control-flow graph, discarding those whose claimed memory operation is unreachable on any feasible path. The validated summaries are injected into CodeQL and Infer via their respective extension mechanisms. Z3 path feasibility filtering then eliminates warnings on infeasible paths, and a final LLM-based validation step confirms whether each remaining warning is a genuine bug. On eight real-world C/C++ projects totaling over 3.6M lines of code, MemHint detects 54 unique memory leaks, all confirmed or fixed, at approximately $1.7 per detected bug, compared to 19 by vanilla CodeQL and 3 by vanilla Infer.

cs.SE↗