arXiv · 2504.04572
Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
Abstract
Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mohamed Eltahir, Osamah Sarraj, Mohammed Bremoo, Mohammed Khurd, Abdulrahman Alfrihidi, Taha Alshatiri, Mohammad Almatrafi, Tanveer Hussain. 2025-04-06. Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric. https://arxiv.org/abs/2504.04572
Cite the original work for its findings. Save a collection to share your selection of sources.