SearcharxivSearch

arXiv subjects

Seyi Olojo

Publications and source records attributed to Seyi Olojo.

5 recordsLinked to original sources

Locating Translation as a Craft in the Age of AI

Rapid development of Large Language Models (LLMs) and similar automated approaches for translation tasks is increasingly affecting the landscape of translation technologies. As concerns about the outsourcing of translator work to these automated translation tools grow, it is increasingly crucial to gather insights from the translation community directly. To this end, we conduct an interview study with 19 professional translators working across 11 languages and 11 domains to understand their perspectives, experiences, and concerns with using translation technologies in their work. We find that translators are cautious when incorporating new tools into their workflow, with several expressing concerns that machine translation (MT) and LLMs are infringing on the necessary human aspects and verification processes of translation. Importantly, translators are worried that these tools have potential for harmful downstream effects due to compromising the human aspects of translation work. These findings demonstrate the need to develop translation technologies that directly serve translators' needs rather than replacing human translation. This can be done by focusing more on the assistive tools that emphasize the uncertain, social, and ultimately human character of translation, rather than automation.

cs.CY

Moral Missions: Surfacing Moral Decision-Making Strategies for Responsible Data Science Practice

A growing ecosystem of techniques, toolkits, and guidelines has been developed to help data scientists consider the social implications of data-driven technologies. However, prior literature highlights that even when this ecosystem of techniques is provided to professional data scientists, they still struggle to consistently adopt a responsible data science practice. We posit that the key to sustained responsible data science practice is to approach it as a moral mission: a conviction-driven technical practice that seeks to transform social conditions by any degree possible. In this paper, we present a semi-structured interview study with 15 responsible data scientists and AI practitioners to understand the moral decision-making procedures they use to articulate and actualize their moral missions. Through a phenomenological analysis of our participants' accounts, we find participants engage in embodied introspection, circumvent institutional expectations, and center relationality throughout their moral missions. We also present how our participants engage in similar processes to contend with generative AI (GenAI) in their responsible practice. We conclude by calling for subversive data science communities and identifying sociotechnical design implications to better support sustainable responsible data science practice.

cs.HC

Shaping the Future of Generative AI for Black Communities: A Frame Analysis of Public Discourse and Empirical Scholarly Research

As generative AI (genAI) systems become embedded in education, employment, healthcare, and creative industries, the impact and engagement among marginalized groups have become both a widespread discourse and a focus in scholarly research. As a starting point, we examine public discourse and empirical research to explore the impact of genAI systems on Black communities. We conducted a systematic literature review (SLR) of 91 empirical papers alongside a media discourse frame analysis of 28 public resources, applying Entman's framing theory to map how each corpus defines problems, attributes causes, and proposes treatments. Our SLR reveals that scholarly research concentrates heavily on technical bias detection, reducing Blackness to measurable variables rather than engaging with cultural practices, structural conditions, or Black knowledge systems. Our frame analysis reveals that public discourse attributes genAI-related harm to historical and systemic forces, while scholarly research stops its causal accounts at the dataset and its treatment recommendations at technical reform. We demonstrate that this misalignment is structurally produced: anti-Blackness operates simultaneously across both registers, generating a shared evacuation of Black epistemic agency. We argue for frame analysis as an AI ethics methodology capable of surfacing what technical evaluation forecloses.

cs.HC

Socially Responsible Data for Large Multilingual Language Models

Large Language Models (LLMs) have rapidly increased in size and apparent capabilities in the last three years, but their training data is largely English text. There is growing interest in multilingual LLMs, and various efforts are striving for models to accommodate languages of communities outside of the Global North, which include many languages that have been historically underrepresented in digital realms. These languages have been coined as "low resource languages" or "long-tail languages", and LLMs performance on these languages is generally poor. While expanding the use of LLMs to more languages may bring many potential benefits, such as assisting cross-community communication and language preservation, great care must be taken to ensure that data collection on these languages is not extractive and that it does not reproduce exploitative practices of the past. Collecting data from languages spoken by previously colonized people, indigenous people, and non-Western languages raises many complex sociopolitical and ethical questions, e.g., around consent, cultural safety, and data sovereignty. Furthermore, linguistic complexity and cultural nuances are often lost in LLMs. This position paper builds on recent scholarship, and our own work, and outlines several relevant social, cultural, and ethical considerations and potential ways to mitigate them through qualitative research, community partnerships, and participatory design approaches. We provide twelve recommendations for consideration when collecting language data on underrepresented language communities outside of the Global North.

cs.CL

The Distressing Ads That Persist: Uncovering The Harms of Targeted Weight-Loss Ads Among Users with Histories of Disordered Eating

Targeted advertising can harm vulnerable groups when it targets individuals' personal and psychological vulnerabilities. We focus on how targeted weight-loss advertisements harm people with histories of disordered eating. We identify three features of targeted advertising that cause harm: the persistence of personal data that can expose vulnerabilities, over-simplifying algorithmic relevancy models, and design patterns encouraging engagement that can facilitate unhealthy behavior. Through a series of semi-structured interviews with individuals with histories of unhealthy body stigma, dieting, and disordered eating, we found that targeted weight-loss ads reinforced low self-esteem and deepened pre-existing anxieties around food and exercise. At the same time, we observed that targeted individuals demonstrated agency and resistance against distressing ads. Drawing on scholarship in postcolonial environmental studies, we use the concept of slow violence to articulate how online targeted advertising inflicts harms that may not be immediately identifiable. CAUTION: This paper includes media that could be triggering, particularly to people with an eating disorder. Please use caution when reading, printing, or disseminating this paper.

cs.HC