TY - RPRT TI - ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research AU - Hao Shen AU - Hang Yang AU - Zhouhong Gu AU - Weili Han PY - 2026 UR - https://arxiv.org/abs/2601.21654 ID - 2601.21654 ER -