arXiv · 2401.13165
Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages
Abstract
This chapter focuses on gender-related errors in machine translation (MT) in the context of low-resource languages. We begin by explaining what low-resource languages are, examining the inseparable social and computational factors that create such linguistic hierarchies. We demonstrate through a case study of our mother tongue Bengali, a global language spoken by almost 300 million people but still classified as low-resource, how gender is assumed and inferred in translations to and from the high(est)-resource English when no such information is provided in source texts. We discuss the postcolonial and societal impacts of such errors leading to linguistic erasure and representational harms, and conclude by discussing potential solutions towards uplifting languages by providing them more agency in MT conversations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sourojit Ghosh, Srishti Chatterjee. 2024-01-24. Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages. https://arxiv.org/abs/2401.13165
Cite the original work for its findings. Save a collection to share your selection of sources.