Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address this limitation, this paper proposes Text-Prompted CLAP (TP-CLAP), a parameter-efficient extension of CLAP that introduces a cross-attention-based fusion module to incorporate textual prompts into audio features. TP-CLAP is trained using an audio multiple-choice question answering (AMCQA) framework, where it learns to align text-conditioned audio representations with text embeddings of correct answer choices via contrastive learning. Experiments demonstrate that TP-CLAP performs competitively with substantially larger audio-LLMs on audio question answering, while also improving the base CLAP model on conventional audio-text retrieval and zero-shot classification benchmarks. The learned representations are further fine-tuned for attribute-focused audio-to-audio retrieval, showing that TP-CLAP consistently outperforms the standard CLAP baseline in music retrieval tasks.