TY - RPRT TI - Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions AU - Mathew Monfort AU - SouYoung Jin AU - Alexander Liu AU - David Harwath AU - Rogerio Feris AU - James Glass AU - Aude Oliva PY - 2021 UR - https://arxiv.org/abs/2105.04489 ID - 2105.04489 ER -