arXiv · 2607.03935
Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions
Abstract
Self-evolving frameworks usually optimize task solutions while treating the surrounding harness as fixed. We introduce Harness-Aware Self-Evolving (HASE), an agentic reinforcement-learning framework in which a single model can generate task solutions or edit selected harness components in a multi-turn action space. HASE enables a single Qwen3-8B model to match the text-classification performance of a GPT-OSS-120B model that uses Claude Code as the harness proposer. In alpha factor mining, HASE outperforms the reported GPT-OSS-120B baseline. HASE also repairs imperfect evaluation components and converges to state-of-the-art performance in circle-packing algorithm discovery. These results show that HASE improves the harness and the solution through one unified agentic process.
Explore related subjects
Keep this discovery
Haochen Luo, Yi Huang, Sichun Luo, Fengyuan Liu, Lei Li, Zefa Hu, Junlan Feng, Qi Liu. 2026-07-04. Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions. https://arxiv.org/abs/2607.03935
Cite the original work for its findings. Save a collection to share your selection of sources.