TY - RPRT TI - JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse AU - Muyao Li AU - Zihao Wang AU - Kaichen He AU - Xiaojian Ma AU - Yitao Liang PY - 2025 UR - https://arxiv.org/abs/2503.16365 ID - 2503.16365 ER -