arXiv · 2608.24048
Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
Abstract
While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.
Explore related subjects
Keep this discovery
Quanwei Tang, Dong Zhang, Shoushan Li, Guodong Zhou. 2026-08-25. Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding. https://arxiv.org/abs/2608.24048
Cite the original work for its findings. Save a collection to share your selection of sources.