arXiv · 2606.24088
Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR
Abstract
Self-supervised learning (SSL) models extract rich speech representations but often come with high-dimensional features, increasing computational complexity. This work explores an SSL-AutoEncoder (SSL-AE) bottlenecking approach to efficiently reduce feature dimensions while maintaining dysarthric Automatic Speech Recognition (ASR) performance. By leveraging an autoencoder, we transform high-dimensional SSL features into a compact space, reducing model complexity and training time. Our method preserves essential speech information, achieving reduced Word Error Rates (WER) while significantly lowering computational costs. Experiments show SSL-AE bottlenecking reduces training time by 8x compared to the SSL baseline, demonstrating efficiency without sacrificing recognition performance. These results highlight AE as an effective solution for SSL feature compression in resource-constrained environments.
Explore related subjects
Keep this discovery
Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Shrikanth Narayanan, Sudarsana Reddy Kadiri. 2026-06-23. Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR. https://arxiv.org/abs/2606.24088
Cite the original work for its findings. Save a collection to share your selection of sources.