Searcharxiv⌕ Search

arXiv · 2610.04779

Asymmetric Repository Lineage Modeling and Verifier-Guided Coordination in Concurrent AI Coding Agents

Abstract

Concurrent AI coding agents create a coordination problem in which cheap signals may prioritize work, but only an executable checker can establish the property being claimed. We study this separation through a property-scoped verification contract and MERGEGYM, a three-track benchmark for open-time scope forecasting, replay- conditioned conflict resolution, and scheduling. On a stratified 715-pair lineage set (167 textual conflicts), 79 conflicts (47.3%) occur despite disjoint authored file sets. This is an operationally important proxy mismatch expected from three-way merge: authored PR diffs are measured against PR-specific bases, whereas the checker compares both heads to their common merge base. A standalone lineage-union rule reaches held-out AUROC 0.877; a 12-feature logistic model reaches 0.882 [0.845, 0.917] with PR-AUC 0.661, and an untuned random forest on the same decision-time features reaches 0.902 [0.875, 0.928] with PR-AUC 0.701. At a 33.3% held-out replay budget, the logistic and random-forest models recover 81.2% and 82.9% of conflicts, respectively. Patch reconstruction succeeds for 48/79 zero- overlap conflicts and all 48 become clean; the other 31 cases are inconclusive, so this check validates the expected three-way-merge explanation rather than claiming a new Git mechanism. In T1, a zero-shot LLM reaches AUC 0.704 and LLM-plus- metadata fusion 0.740. In T3, a decision-time gate de-overlaps a median 91.7% of labeled scope collisions at 65.0% makespan inflation under frozen-label replay. Because local git merge-tree replay is already cheap in our logs (median 0.02 s), we do not claim that lineage triage saves this checker alone: when exact replay is cheap, verify everything. All empirical guarantees in this paper remain limited to textual mergeability or the explicitly stated frozen-label scheduling target.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Arjun Subramanian, George Xu, Nithilan Karthik. 2026-10-03. Asymmetric Repository Lineage Modeling and Verifier-Guided Coordination in Concurrent AI Coding Agents. https://arxiv.org/abs/2610.04779

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Can we find bugs using LLM-generated oracles?

Unit testing is vital in software development. Typically, a unit test consists of a test prefix and a test oracle which captures the developer's intended behaviour. Traditional test generation tools (e.g. Randoop and Evosuite) often produce oracles that mirror the program's actual behavior rather than the expected one, limiting their ability to automatically detect bugs as users must manually verify if the generated assertions are correct. Recent approaches leverage Large Language Models (LLMs), trained on vast datasets, to generate developer-like code and test cases. Although successful in generating tests, the question of whether such LLM-generated oracles can automatically find bugs, i.e., expected software behavior, remains unanswered. We conduct a controlled experiment to answer this question, by studying LLMs on two tasks, namely, test oracle classification and generation, and assessing whether LLM oracles capture the actual or the expected behavior. The study includes test cases and oracles written by developers and automatically generated for 24 Java repositories. Our findings show that LLM-based test generation approaches mainly capture the actual program behavior making bug detection difficult. We also find that LLMs are better at generating oracles than classifying them. Notably, LLM-generated oracles have a higher fault detection potential than the Evosuite ones.

cs.SE↗

SELU: A Software Engineering Language Understanding Benchmark

Large Language Models (LLMs) have demonstrated remarkable capabilities in code understanding and generation. However, their effectiveness on non-code Software Engineering (SE) tasks remains underexplored. We present 'Software Engineering Language Understanding' (SELU), the first comprehensive benchmark for evaluating LLMs on 22 SE textual artifacts NLU tasks, spanning from identifying whether a requirement is functional or non-functional to estimating the effort required to implement a development task. SELU covers classification, regression, Named Entity Recognition (NER), and Masked Language Modeling (MLM) tasks, with data drawn from diverse sources such as issue tracking systems and developer forums. We fine-tune 22 open-source LLMs, both generalist and domain-adapted; and prompt two proprietary alternatives using zero-shot a 3-shot prompting strategies. Performance is measured using metrics such as F1-macro, SMAPE, F1-micro, and accuracy, and compared via the Bayesian signed-rank test. Our results show that fine-tuned models across various sizes and architectures perform best, exhibiting high mean performance and low across-task variance. Furthermore, domain adaptation via code-focused pre-training does not yield significant improvements and might even be counterproductive for developer communication tasks.

cs.SE↗

Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports

Mutation-based fuzzing is effective for uncovering compiler bugs, but designing high-quality mutators for modern languages with complex constructs (e.g., templates, macros) remains challenging. Existing methods rely heavily on manual design or human-in-the-loop correction, limiting scalability and cross-language generalizability. We present Mut4All, a fully automated, language-agnostic framework that synthesizes mutators using Large Language Models (LLMs) and compiler-specific knowledge from bug reports. It consists of three agents: (1) a mutator invention agent that identifies mutation targets and generates mutator metadata using compiler-related insights; (2) a mutator implementation synthesis agent, fine-tuned to produce initial implementations; and (3) a mutator refinement agent that verifies and corrects the mutators via unit-test feedback. Mut4All processes 1400 bug reports (700 Rust, 700 C++), yielding 444 Rust and 561 C++ mutators at ~$0.08 each via GPT-4o. Our customized fuzzer, using these mutators, finds 62 bugs in Rust compilers (44 new, 32 fixed) and 38 bugs in C++ compilers (17 new, 3 fixed). Mut4All outperforms existing methods in both unique crash detection and coverage, ranking first on Rust and second on C++.

cs.SE↗