arXiv · 2511.07595
TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
Also available from
Abstract
In this work, we introduce TurkEmbed4Retrieval, a retrieval specialized variant of the TurkEmbed model originally designed for Natural Language Inference (NLI) and Semantic Textual Similarity (STS) tasks. By fine-tuning the base model on the MS MARCO TR dataset using advanced training techniques, including Matryoshka representation learning and a tailored multiple negatives ranking loss, we achieve SOTA performance for Turkish retrieval tasks. Extensive experiments demonstrate that our model outperforms Turkish colBERT by 19,26% on key retrieval metrics for the Scifact TR dataset, thereby establishing a new benchmark for Turkish information retrieval.
Explore related subjects
Keep this discovery
Özay Ezerceli, Gizem Gümüşçekiçci, Tuğba Erkoç, Berke Özenç. 2025-11-10. TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task. https://doi.org/10.1109/asyu67174.2025.11208511
Cite the original work for its findings. Save a collection to share your selection of sources.