TY - RPRT TI - Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization AU - Xuefeng Liu AU - Mingxuan Cao AU - Qinan Huang AU - Thomas Brettin AU - Rick Stevens AU - Le Cong PY - 2026 UR - https://arxiv.org/abs/2607.00531 ID - 2607.00531 ER -