arXiv · 2606.21658
Towards LLM-Powered Automation of a Dark Matter Constraint Repository
Abstract
Dark matter constraint repositories are critical community infrastructure, yet the most widely-used are volunteer-maintained, a sustainability risk as results accelerate. Limits are typically published as figures rather than machine-readable records, compounding the burden. We present a large language model (LLM) pipeline that monitors arXiv, extracts limit curves, integrates them as code, and opens pull requests for human review. On a 329-paper benchmark of measured limits graded against the upstream-curated repository, it identifies the coupling type for 98.1% of gradable papers. The 81% yielding a comparable extraction reach a median residual of 0.17 dex from a single model read, and 51% of all papers land within a factor of two of the curator's curve. The LLM is confined to perception; arithmetic, unit conversion, and the guards that reject mis-extractions are deterministic, testable code. Deployed, the pipeline has generated limit proposals, but none have merged: governance of AI-generated scientific data is itself unsolved.
Explore related subjects
Keep this discovery
Lanqing Yuan, Karthik Ramanathan. 2026-06-19. Towards LLM-Powered Automation of a Dark Matter Constraint Repository. https://arxiv.org/abs/2606.21658
Cite the original work for its findings. Save a collection to share your selection of sources.