TY - RPRT TI - Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model AU - Hung-Chieh Fang AU - Nai-Xuan Ye AU - Yi-Jen Shih AU - Puyuan Peng AU - Hsuan-Fu Wang AU - Layne Berry AU - Hung-yi Lee AU - David Harwath PY - 2024 UR - https://arxiv.org/abs/2402.05819 ID - 2402.05819 ER -