arXiv · 2608.19523
DAVSS: Distilled Audio-Visual State Space Models
Abstract
State-space models (SSMs) distilled from transformer teachers combine the performance of transformers with the efficiency of SSMs. We extend the Transformer-SSM knowledge distillation to a multimodal setting and propose the Distilled Audio-visual State-Space (DAVSS) model. The DAVSS model, 14M parameters, is 12 times smaller compared to transformer-based models such as CAV-MAE, and still outperforms them. DAVSS improves over the existing audio-visual models by: 1) Finer input resolution: using smaller patch sizes process the input, compensating for the smaller model size by increasing input sequence lengths. This is supported by the observation that a larger patch size results in lower performance. 2) Deeper joint modeling: utilizing a larger portion of the model (30%) for joint audio-visual processing, compared to <5% in CAV-MAE, enabling deeper cross-modal interaction without significantly increasing the computational cost associated with the concatenated audio-visual tokens.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Saurabhchand Bhati, Mrudula Athi, Amit S. Chhetri, James Glass. 2026-08-20. DAVSS: Distilled Audio-Visual State Space Models. https://arxiv.org/abs/2608.19523
Cite the original work for its findings. Save a collection to share your selection of sources.