TY - RPRT TI - CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models AU - Hao-Wen Dong AU - Xiaoyu Liu AU - Jordi Pons AU - Gautam Bhattacharya AU - Santiago Pascual AU - Joan SerrĂ  AU - Taylor Berg-Kirkpatrick AU - Julian McAuley PY - 2023 UR - https://arxiv.org/abs/2306.09635 ID - 2306.09635 ER -