TY - RPRT TI - Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos AU - Daniele Materia AU - Francesco Ragusa AU - Giovanni Maria Farinella PY - 2026 UR - https://arxiv.org/abs/2604.03667 ID - 2604.03667 ER -