arXiv · 2609.37426
LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension
Abstract
Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs--Gemma 4 31B and Qwen3.6 27B--across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matarić. 2026-09-29. LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension. https://arxiv.org/abs/2609.37426
Cite the original work for its findings. Save a collection to share your selection of sources.