arXiv · 2310.03724
Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer
Abstract
Recent research has shown that independently trained encoders and decoders, combined through a shared fixed-size representation, can achieve competitive performance in speech-to-text translation. In this work, we show that this type of approach can be further improved with multilingual training. We observe significant improvements in zero-shot cross-modal speech translation, even outperforming a supervised approach based on XLSR for several languages.
Explore related subjects
Keep this discovery
Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot. 2023-10-05. Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer. https://doi.org/10.21437/interspeech.2023-2484
Cite the original work for its findings. Save a collection to share your selection of sources.