TY - RPRT TI - ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation AU - Ali Athar AU - Xueqing Deng AU - Liang-Chieh Chen PY - 2025 UR - https://arxiv.org/abs/2412.09754 ID - 2412.09754 ER -