TY - RPRT TI - 3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks AU - Vineet Bhat AU - Yu-Hsiang Lan AU - Prashanth Krishnamurthy AU - Ramesh Karri AU - Farshad Khorrami PY - 2026 UR - https://arxiv.org/abs/2505.05800 ID - 2505.05800 ER -