TY - RPRT TI - Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens AU - Minsu Kim AU - Jeongsoo Choi AU - Soumi Maiti AU - Jeong Hun Yeo AU - Shinji Watanabe AU - Yong Man Ro PY - 2023 UR - https://arxiv.org/abs/2309.08531 ID - 2309.08531 ER -