arXiv · 2609.39640
Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
Abstract
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zarè Palanciyan, Joaquin Vanschoren. 2026-09-30. Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies. https://arxiv.org/abs/2609.39640
Cite the original work for its findings. Save a collection to share your selection of sources.