SearcharxivSearch

arXiv · 2607.02774

Was It Never Collected, or Rewritten Away? A Commit-Provenance Dataset Separating Ingestion Gaps from Upstream History Edits across the World of Code

Abstract

Any global mirror of open-source version control is incomplete, but the reasons a commit is missing are not interchangeable: a project may have force-pushed it away (so it no longer exists upstream), or the mirror may never have ingested it (a true collection gap). We release a commit-provenance dataset that separates these cases at scale by comparing two views of the same commit graph: the GHArchive event stream, a historical witness of what GitHub advertised at push time, and the World of Code (WoC) V2604 object database, an accumulated union of periodic fetches that never deletes a collected commit. Walking each reference's PushEvent chain reconstructs force-pushes structurally (a before that is not the prior head breaks the fast-forward chain; no recorded flag exists across eras), and joining every advertised commit against WoC membership yields a three-way label. Over 1,118,116,350 advertised commits, 53.35% are present in WoC, 6.47% are rewritten (orphaned by a later force-push, an upstream edit and a correct absence), and 40.18% are never-ingested (the candidate collection gap). About one missing-commit case in fifteen is a rewrite the project erased, not a mirror gap. We release the force-push witness (166,710,831 events over 19,926,250 repositories) and a per-project rollup (78,125,788 repositories; force-pushing observed in 25.47%, a rewritten commit in 12.85%). We treat the 40% never-ingested share as an upper bound on the collection gap and state five reasons it overcounts. The dataset makes mirror-completeness reporting honest, flags rewritten commits as duplicate patch sources for contribution counting, and corrects a 10.82% corpus-wide undercount that history editing imposes on commit-based productivity, with a released per-project correction factor. All artifacts ship as a self-contained, independently hosted replication package keyed to WoC V2604.

Explore related subjects

Keep this discovery

BibTeXRIS

Audris Mockus. 2026-07-02. Was It Never Collected, or Rewritten Away? A Commit-Provenance Dataset Separating Ingestion Gaps from Upstream History Edits across the World of Code. https://arxiv.org/abs/2607.02774

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories

Large language models (LLMs) are increasingly used to generate and assist with software development, yet existing vulnerability datasets largely focus on human-written code or controlled prompting environments. This limits the ability to study security weaknesses in LLM-generated code as it appears in real-world software projects. We present LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories. We mine AI-assisted development activity from GitHub over a 4 year period, from November 13, 2022 to September 3, 2026, using provenance signals such as commit metadata and AI-related authorship evidence. After filtering and deduplication, LLMVul contains 21,430 unique C/C++ functions from 226 repositories, together with repository, commit, function, provenance, and AI-tool metadata. We establish vulnerability labels using an ensemble of complementary static-analysis and pattern-based techniques and assign Common Weakness Enumeration (CWE) categories to confirmed vulnerable functions. To assess labeling reliability, we additionally conduct independent manual annotation and measure inter-rater agreement using Cohen's kappa ($k=0.79$). LLMVul contains 1,540 ensemble-vulnerable functions spanning 17 unique CWE categories, providing substantially more real-world LLM-generated vulnerable C/C++ functions than existing vulnerability-oriented LLM code benchmarks. By preserving both code-level vulnerability labels and generation/provenance metadata, LLMVul enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development. The LLMVul dataset is publicly available at https://doi.org/10.5281/zenodo.22668216.

cs.SE

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

Studies of the Model Context Protocol (MCP) server ecosystem draw their samples in ways that quietly select for servers that work: reference sets, popularity lists, hand-curated frames, or pipelines that repair a server until it starts. We report what an unrepaired probability sample actually contains. From a 24,135-server registry census we draw 400 npm/stdio servers with a published seed and probe each one over the wire. Only 48.8% complete an initialize handshake, against 66.7% for a hand-curated frame measured with the same instrument, and the dominant failure is not missing credentials (13.3%) but servers that never start at all (37.5%). Among the 195 that do run, hard conformance is total: zero fatal JSON Schema violations across 2,766 advertised tools. Optional safety annotations are the real variance, and the tool-level omission rate on a random draw is 58.8% against 41.5% on the curated frame, so curation flatters this figure too. We then compare the tool descriptions these servers advertise against two tool-use benchmark corpora using one method held constant. Real MCP tools show 2.8% near-duplication at cosine 0.70, and all of it lies within single servers: cross-author near-duplication is 0.0% at every threshold tested. BFCL v4 shows 16.7%, of which 16.4 points lie between independently presented tasks. UltraTool shows 0.3%, cleaner than real tools, so this is a property of BFCL and not of synthetic corpora as a class. Separately, 68.8% of raw BFCL rows and 85.6% of raw UltraTool rows are exact name-plus-description repeats, against 0.4% for real MCP, so any statistic computed over these releases without global deduplication measures repetition rather than tools. All figures regenerate from released scripts and a published seed.

cs.SE

Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures

Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic benchmark for cost-aware commit gates. Its 48 task templates yield 2,880 scenarios across six fault regimes. A fixed-call 2 x 2 experiment separates verifier-model diversity from evidence-source diversity. On frozen proposals from two local actor families, a cross-model vote over shared evidence approves 62.9% of unsafe proposals, versus 22.9% with an independent source. The source effect is 40.9 percentage points, compared with 11.3 for model diversity. A portfolio controller selects verification plans using only deployment-observable metadata. Approximate cluster-adjusted calibration at a nominal 5% per-task target yields 1.9% unsafe execution and 38.2% automated safe coverage on the locked test. Matched-budget portfolios also improve on fixed verification policies. Transfer remains conditional: unseen fault families yield 16-26% risk, and a FinQA check fails to reproduce the source effect with the tested small verifiers. A preregistered live HTTP/SQLite study tests concurrent writes and lost responses. After-check races defeat verifier-only gates; transactional partial guards prevent only covered failures, while a full atomic guard records no unsafe effects across 216 episodes. Idempotent request identifiers prevent duplicate effects after lost responses. The results motivate explicit evidence lineage, cost-aware selection, and commit-time enforcement, while exposing the limits of approximate calibration and local-tool generalization.

cs.SE