arXiv · 2409.18952
RepairBench: Leaderboard of Frontier Models for Program Repair
Abstract
AI-driven program repair uses AI models to repair buggy software by producing patches. Rapid advancements in AI surely impact state-of-the-art performance of program repair. Yet, grasping this progress requires frequent and standardized evaluations. We propose RepairBench, a novel leaderboard for AI-driven program repair. The key characteristics of RepairBench are: 1) it is execution-based: all patches are compiled and executed against a test suite, 2) it assesses frontier models in a frequent and standardized way. RepairBench leverages two high-quality benchmarks, Defects4J and GitBug-Java, to evaluate frontier models against real-world program repair tasks. We publicly release the evaluation framework of RepairBench. We will update the leaderboard as new frontier models are released.
Explore related subjects
Keep this discovery
André Silva, Martin Monperrus. 2024-09-27. RepairBench: Leaderboard of Frontier Models for Program Repair. https://doi.org/10.1109/llm4code66737.2025.00006
Cite the original work for its findings. Save a collection to share your selection of sources.