arXiv · 2609.16495
Segmental Posterior Decoding for Audio Moment Retrieval
Abstract
Audio moment retrieval (AMR) identifies temporal segments in long recordings that best match a free-form text query. Existing systems largely rely on fixed-slot DETR decoders that assign proposal-level confidence scores without explicitly normalizing over competing explanations of the full timeline. We propose segmental posterior decoding, which defines a globally normalized distribution over temporal segmentations and scores each candidate moment by its exact segment marginal posterior computed through forward-backward inference. We further expand the training segmentation space by treating a foreground span and its adjacent subdivisions as distinct hypotheses, thereby increasing competition among alternative segmentations. On CASTELLA, our method achieves 41.15% R1@0.7 and 34.68% mAP, outperforming the same network decoded with DETR slot confidence by 10.91 and 9.20 percentage points, respectively.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seungdeok Choi, Seongmin Choi, Inhan Choi, Junho Kim, Jeong-gyu Ban, Yong-Hwa Park. 2026-09-15. Segmental Posterior Decoding for Audio Moment Retrieval. https://arxiv.org/abs/2609.16495
Cite the original work for its findings. Save a collection to share your selection of sources.