arXiv · 2503.23671
CrossFormer: Cross-Segment Semantic Fusion for Document Segmentation
Abstract
Text semantic segmentation involves partitioning a document into multiple paragraphs with continuous semantics based on the subject matter, contextual information, and document structure. Traditional approaches have typically relied on preprocessing documents into segments to address input length constraints, resulting in the loss of critical semantic information across segments. To address this, we present CrossFormer, a transformer-based model featuring a novel cross-segment fusion module that dynamically models latent semantic dependencies across document segments, substantially elevating segmentation accuracy. Additionally, CrossFormer can replace rule-based chunk methods within the Retrieval-Augmented Generation (RAG) system, producing more semantically coherent chunks that enhance its efficacy. Comprehensive evaluations confirm CrossFormer's state-of-the-art performance on public text semantic segmentation datasets, alongside considerable gains on RAG benchmarks.
Explore related subjects
Keep this discovery
Tongke Ni, Yang Fan, Junru Zhou, Xiangping Wu, Qingcai Chen. 2025-03-31. CrossFormer: Cross-Segment Semantic Fusion for Document Segmentation. https://arxiv.org/abs/2503.23671
Cite the original work for its findings. Save a collection to share your selection of sources.