Searcharxiv⌕ Search

arXiv · 2610.03675

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Abstract

LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the resulting code. We also design a cache-efficient evolution process, where our harness and prompts maximize the sharing of prefixes across different evolution steps, to improve cache reuse. To measure solution quality throughout a fixed cost budget, we introduce Budget-Aware Area Under the Curve (BA-AUC), defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget. Across 10 mathematical and systems optimization tasks, FrugalEvo matches or surpasses state-of-the-art baselines, including OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX, in final solution quality and achieves higher BA-AUC on 9 tasks. It also achieves higher average performance than these baselines on 10 algorithmic optimization tasks from ALE-Bench-Lite. Notably, on circle packing, FrugalEvo achieves new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD and with GLM-5.3 and its Flash variant for only 0.55 USD, matching or surpassing all baselines, including multi-agent methods such as CORAL and SwarmResearch, which cost approximately 50 USD on average.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hui Chen, Xuan Qi, James Xu Zhao, Zhaopeng Feng, Shilong Liu, Kuang Xu, Pang Wei Koh, Bryan Hooi. 2026-10-02. FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution. https://arxiv.org/abs/2610.03675

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets

Spike-based encodings are sparse and energy-efficient. However, previous theory for such spiking representations is largely disconnected from most signal processing literature and does not extend to multi-scale frames. We recast spike-encoders as time-causal overcomplete wavelet frames with closed-form bandwidth tiling and reconstruction error bounds. We show that scale covariance, frame bounds, and reconstruction guarantees hold for any finite-energy input in both encoding and decoding steps. The proposed spiking bandpass wavelets preserve the sparsity and locality of spiking representations: A threshold parameter sets the event rate, and the reconstruction error grows linearly with it, up to spike quantization and time discretization. On ECG and audio datasets, we reach reconstruction accuracy comparable to continuous wavelet transforms, multi-channel integrate-and-fire time encoding machines, and Sigma-Delta modulation. The experiments demonstrate robustness of the proposed spiking bandpass wavelets to noise and spike-time jitter, and that the spike activity shifts across channels as predicted by the covariance property under temporal rescaling. Every operation reduces to a leaky integrate-and-fire difference with thresholding, and both the encoder and the decoder map directly to existing neuromorphic hardware.

cs.NE↗

Continual Reinforcement Learning with Neuroevolution

Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.

cs.NE↗

Evolutionary Computation for Trustworthy AI: From Attacks and Defenses to Self-Evolving Era

As Artificial Intelligence (AI) has evolved from task-specific models to foundation models and agents, the scope of trustworthy AI has expanded from model-level robustness to the reliability and safety of broader AI systems. This evolution has also expanded the attack surface from individual models to broader system-level interactions, including tool use, context, and interaction trajectories with dynamic environments. As a result, maintaining reliable and safe behavior under changing or deliberately manipulated conditions has become increasingly challenging. The search for effective attacks and defenses often relies on black-box feedback to navigate discrete choices among words, actions, system components, or their combinations. Multiple objectives and expensive candidate evaluations further limit what can be explored. Evolutionary Computation (EC), with its population-based, gradient-free search and flexible variation and selection mechanisms, is well suited to these settings. This survey reviews how EC has been applied to trustworthy AI across three directions: evolutionary attacks, evolutionary defenses, and trustworthy self-evolving AI systems. Unlike prior reviews that treat trustworthy AI, EC, and self-evolving systems largely separately, we connect these lines through a common evolutionary perspective. For self-evolving AI, we examine how trustworthiness governs the generation and retention of updates that shape subsequent adaptation. We further synthesize evaluation methods and benchmark resources from both trustworthiness and evolutionary-search perspectives. Finally, we discuss key challenges and future research directions toward more effective and reliable use of EC in trustworthy AI.

cs.NE↗