arXiv · 2509.05703
Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis
Abstract
Marine mammal vocalization analysis depends on interpreting bioacoustic spectrograms. Vision Language Models (VLMs) are not trained on these domain-specific visualizations. We investigate whether VLMs can extract meaningful patterns from spectrograms visually. Our framework integrates VLM interpretation with LLM-based validation to build domain knowledge. This enables adaptation to acoustic data without manual annotation or model retraining.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ragib Amin Nihal, Benjamin Yen, Takeshi Ashizawa, Kazuhiro Nakadai. 2025-09-06. Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis. https://arxiv.org/abs/2509.05703
Cite the original work for its findings. Save a collection to share your selection of sources.