TY - RPRT TI - Align before Fuse: Vision and Language Representation Learning with Momentum Distillation AU - Junnan Li AU - Ramprasaath R. Selvaraju AU - Akhilesh Deepak Gotmare AU - Shafiq Joty AU - Caiming Xiong AU - Steven Hoi PY - 2021 UR - https://arxiv.org/abs/2107.07651 ID - 2107.07651 ER -