arXiv · 2601.15236
MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering
Abstract
Large language models can memorize competing locale-specific facts yet fail to select among them when the locale changes, defaulting instead to a single globally dominant answer. We formalize this as localized knowledge disambiguation and introduce LocalNewsQA, an 18,700-item English-news benchmark that pairs the same question across two locales and scores whether a model actually switches its answer when the locale changes. We also introduce MAPLE, a controlled family of decoder-only models pretrained with document-level geographic metadata (source URL, country, and continent) already present in the training corpus, and compare it to metadata-free controls trained on identical data with the same token budget, architecture, and optimization. In controlled experiments at 1B and 3B, with inference-time metadata fixed, pretraining with metadata in MAPLE produces measurable switching and improves accuracy on questions whose correct answer depends on locale. Ablations and external-benchmark evaluations further suggest that locale-conditioned prediction benefits from geographic provenance learned during pretraining and that these benefits strengthen at larger model sizes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos. 2026-01-21. MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering. https://arxiv.org/abs/2601.15236
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.