arXiv · 2512.11388
Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis
Abstract
We investigated the impact of data selection on machine translation fine-tuning for open LLMs. Using Japanese-English corpora, we compare five selectors: TF-IDF, COMET Kiwi, QuRate, FD-Score, and random selection, under controlled training conditions. We observed that semantic selectors consistently outperform lexical and geometry-based heuristics, and that even when the selected data differ by less than 3%, the impact on model performance is substantial, underscoring the sensitivity of fine-tuning to data quality.
Explore related subjects
Keep this discovery
Felipe Ribeiro Fujita de Mello, Hideyuki Takada. 2025-12-12. Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis. https://doi.org/10.1109/bigdata66926.2025.11402145
Cite the original work for its findings. Save a collection to share your selection of sources.