TY - RPRT TI - Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation AU - Shih-Lun Wu AU - Xuankai Chang AU - Gordon Wichern AU - Jee-weon Jung AU - François Germain AU - Jonathan Le Roux AU - Shinji Watanabe PY - 2024 UR - https://arxiv.org/abs/2309.17352 ID - 2309.17352 ER -