arXiv · 2510.12326
DeePAQ: A Perceptual Audio Quality Metric Based On Foundational Models and Weakly Supervised Learning
Abstract
This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to construct an embedding space that captures distortion intensity in general audio. To the best of our knowledge, DeePAQ is the first in the general audio quality domain to leverage weakly supervised labels and metric learning for fine-tuning a music foundation model with Low-Rank Adaptation (LoRA), a direction not yet explored by other state-of-the-art methods. We benchmark the proposed model against state-of-the-art objective audio quality metrics across listening tests spanning audio coding and source separation. Results show that our method surpasses existing metrics in detecting coding artifacts and generalizes well to unseen distortions such as source separation, highlighting its robustness and versatility.
Explore related subjects
Keep this discovery
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado, Jürgen Herre. 2025-10-14. DeePAQ: A Perceptual Audio Quality Metric Based On Foundational Models and Weakly Supervised Learning. https://arxiv.org/abs/2510.12326
Cite the original work for its findings. Save a collection to share your selection of sources.