TY - RPRT TI - CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment AU - Hongwei Xue AU - Yuchong Sun AU - Bei Liu AU - Jianlong Fu AU - Ruihua Song AU - Houqiang Li AU - Jiebo Luo PY - 2023 UR - https://arxiv.org/abs/2209.06430 ID - 2209.06430 ER -