TY - RPRT TI - An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction AU - Guanting Shen AU - Zi Tian PY - 2026 UR - https://arxiv.org/abs/2602.20219 ID - 2602.20219 ER -