arXiv · 2510.19585
Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
Abstract
This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin.
Explore related subjects
Keep this discovery
Yu Wu, Ke Shu, Jonas Fischer, Lidia Pivovarova, David Rosson, Eetu Mäkelä, Mikko Tolonen. 2025-10-22. Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark. https://arxiv.org/abs/2510.19585
Cite the original work for its findings. Save a collection to share your selection of sources.