arXiv · 2610.08300
Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval
Abstract
Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural depth and the width of context supplied to the answer model. Using EverMemBench, we evaluate depths D1-D4, core budgets of 1,024/2,048/4,096 tokens, and additional Production and Oracle conditions up to the full archive. Increasing width from 1K to 4K improves Accuracy by 10.11-17.98 percentage points, whereas increasing depth provides no monotonic gain. Beyond 8-16K, Production performance reaches a plateau while tokens per correct answer continue to increase; Oracle preserves quality on full archives of 68-71K tokens. These results motivate further investigation of large, coherent context blocks instead of progressively deeper memory structures.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Andreev. 2026-10-06. Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval. https://arxiv.org/abs/2610.08300
Cite the original work for its findings. Save a collection to share your selection of sources.