arXiv · 2607.00596
Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs
Abstract
This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric heuristics, YOLO-based layout parsing, an end-to-end document model ECLAIR, and a hybrid method combining semantic zone detection with a generative LLM. Our hybrid method achieves the lowest error rates of all evaluated approaches, reducing ordering errors by up to 76% over the strongest geometric baseline, and remains robust in multi-page settings and under noisy OCR. Rather than targeting production the method is designed as a data bootstrapping strategy enabling rapid annotation in highly under-resourced scenarios. Alongside the dataset, we release a specialized Tesseract OCR model for historical Armenian print.
Explore related subjects
Keep this discovery
Chahan Vidal-Gorène, Nadi Tomeh, Victoria Khurshudyan. 2026-07-01. Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs. https://arxiv.org/abs/2607.00596
Cite the original work for its findings. Save a collection to share your selection of sources.