arXiv · 2509.14689
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
Abstract
Large pre-trained speech models excel in downstream tasks but their deployment is impractical for resource-limited environments. In this paper, we introduce HArnESS, the first Arabic-centric self-supervised speech model family, designed to capture Arabic speech nuances. Using iterative self-distillation, we train large bilingual HArnESS (HL) SSL models and then distill knowledge into compressed student models (HS, HST), preserving Arabic-specific representations. We use low-rank approximation to further compact the teacher's discrete supervision into shallow, thin models. We evaluate HArnESS on Arabic ASR, Speaker Emotion Recognition (SER), and Dialect Identification (DID), demonstrating effectiveness against HuBERT and XLS-R. With minimal fine-tuning, HArnESS achieves SOTA or comparable performance, making it a lightweight yet powerful alternative for real-world use. We release our distilled models and findings to support responsible research and deployment in low-resource settings.
Explore related subjects
Keep this discovery
Vrunda N. sukhadia, Shammur Absar Chowdhury. 2025-09-18. HARNESS: Lightweight Distilled Arabic Speech Foundation Models. https://arxiv.org/abs/2509.14689
Cite the original work for its findings. Save a collection to share your selection of sources.