arXiv · 2608.21237
Indexing Long Documents for LLM-Based Analysis
Abstract
Long documents such as clinical records, legal contracts, and scientific papers are increasingly analyzed with large language models (LLMs). Naturally, feeding the full document to the model for every question can eventually become slow, expensive, prone to hallucination, and it reuses no work across questions. We explore an indexing-based solution for document analysis and propose a hierarchical plain-text index that is built once per document and consulted by subsequent queries. Inspired by the classic B+ tree, the index organizes a document into pages arranged from general summaries at the root to specific ones at the leaves, but it departs from the B+ tree in three ways suited to LLM access: content lives at every level, the index is plain text rather than attribute values, and navigation follows relevance between page summaries rather than comparing a search key. Its structure is discovered per document by the LLM rather than being hand-designed. In a preliminary evaluation on the NarrativeQA dataset, the index reaches accuracy comparable to DocETL, the strongest baseline, while being 40\% cheaper to answer questions.
Explore related subjects
Keep this discovery
Donna Pham. 2026-08-21. Indexing Long Documents for LLM-Based Analysis. https://arxiv.org/abs/2608.21237
Cite the original work for its findings. Save a collection to share your selection of sources.