arXiv · 2609.32227
OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?
Abstract
Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets. The testbed uses two optimization regimes, surface obfuscation controls, calibrated references, held-out/stress splits, and diagnostics for degradation and exceptional failures, with LLM API cost reported separately from local evaluator wall-clock. The empirical study asks three questions: whether models can close the calibrated gap between a designated weak starter and an editable competent baseline, whether they can refine editable competent baselines without damaging them, and whether gains survive surface obfuscation controls. Across twelve frontier LLMs and five games, models improve designated weak starters more consistently than they refine editable competent baselines, with substantial variation across games and models. OptiArena provides a practical testbed for measuring bounded-resource algorithm optimization within the five-edit, fixed-scaffold setting studied here. Code is available at https://github.com/WJ-Peng/OptiArena.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenjun Peng, Xinyu Wang. 2026-09-26. OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?. https://arxiv.org/abs/2609.32227
Cite the original work for its findings. Save a collection to share your selection of sources.