arXiv · 2410.04422
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
Abstract
Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well addressed with a sufficient number of reasoning steps, guided by specific CoT prompts. This result emphasizes the potential necessity of solving specific long-context tasks using long-CoT methods, while previous long-context benchmarks always ignore the necessity of long reasoning for long-context tasks and treat them as direct QA tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yijiong Yu, Yongfeng Huang, Zhixiao Qi, Wei Wang, Weifeng Liu, Ran Chen, Ji Pei. 2024-10-06. Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps. https://arxiv.org/abs/2410.04422
Cite the original work for its findings. Save a collection to share your selection of sources.