Searcharxiv⌕ Search

arXiv subjects

Minh Duc Chu

Publications and source records attributed to Minh Duc Chu.

11 recordsLinked to original sources

When Chatbots Accommodate: Auditing the Response Policies of AI Companions in Vulnerable Conversations

Millions turn to AI companion chatbots during loneliness, grief, and personal crises. How these companion platforms respond in such moments can shape the trajectory of a user's vulnerable state. Yet existing model audits evaluate reactions to pre-defined crisis prompts and miss the response policy that governs sustained real-world interaction. We address these gaps with two key contributions. First, we introduce the AI Companion Vulnerability-Response Taxonomy, a grounded, paired taxonomy of user vulnerability and chatbot response designed for analyzing extended companion chatbot interactions. Second, we apply Maximum Causal Entropy Inverse Reinforcement Learning to ~47k turns of real-world user conversations with GPT-4.1, Character.AI, and Replika to infer each platform's short-horizon response policy: the probability of each response category given the user's current vulnerability state. Our findings reveal distinct response profiles of AI companions in conversations with vulnerable users: GPT-4.1 reaches for advice, Character.AI spreads its response across different strategies, and Replika consistently asks questions and stays present. Over four weeks of repeated interaction, GPT-4.1 asks progressively fewer follow-up questions when users are distressed and increasingly sets boundaries or refers users out rather than pushing back. Within each platform, exploratory comparisons across user groups suggest that response policies also differ with users' pre-existing psychological risks and their bonds with the companion. Estimated model response policies are invisible to shallow behavioral audits, providing a new lens for auditing chatbots in the wild and enabling more realistic safety evaluation.

cs.HC↗

The Body as Status: Muscularity, Engagement, and Body Image Risk on #GymTok

Body image concerns among boys and young men are increasingly oriented toward muscularity, with social media serving as a central context for communicating and evaluating these ideals. While prior research has focused on the thin-ideal, less is known about how the muscular-ideal is represented and reinforced on visual social media platforms. This study examines (1) dominant content themes, (2) perceived harm to body image, and (3) engagement patterns across #GymTok, a muscularity-oriented fitness subculture on TikTok. We conducted a content analysis of 2,210 #GymTok videos annotated by clinical experts across themes like self-objectification, rigid dieting, excessive exercise, supplement and steroid use, and masculinity. Annotators also rated the perceived harm of videos to the viewers' body image, and depicted bodies were coded according to muscularity level. Perceived harm varied across content themes, with supplement- and steroid-related content rated as most harmful. Engagement was positively associated with both muscularity and perceived harm: videos depicting more muscular bodies and those rated as more harmful received greater views, likes, shares, and comments. Although less prevalent, masculinity-focused content generated the highest engagement. These findings suggest that TikTok may not only expose users to muscular ideals and potentially harmful behaviors, but also algorithmically amplify them. By increasing the visibility of highly muscular and harmful content, recommendation systems may intensify social comparison processes, while objectification elevates the muscular body into a marker of status, masculinity, and social worth. Together, these dynamics may contribute to body image risk among boys and young men.

cs.CY↗

Tied In on TikTok: Tie Strength and Emotional Dynamics in Algorithmic Communities

Whether genuine communities can form on algorithmically-driven short-form video platforms like TikTok remains an open question, given that user interactions are often brief, dispersed, and difficult to trace. Building on theories of tie strength and online community formation, we examine whether eating disorder (ED) discourse on TikTok exhibits behavioral and emotional signatures of strong ties, including more frequent, reciprocal, and affectively intense interactions. In this paper, we analyze 43,040 ED-related TikTok videos and over 560,000 comments, alongside a Non-ED comparison dataset. We find that at the user-pair level, greater interaction frequency is associated with increasingly positive emotional expression, a pattern that is amplified in ED-related conversations. This trend is also reflected linguistically, with pairs that interact more frequently exhibiting more of a positive tone. At the same time, how a relationship starts matters: pairs that begin with positive exchanges usually stay mostly positive as they continue interacting, while pairs that begin negatively may add some positive exchanges over time but rarely become mostly positive. To contextualize these dynamics, we classify ED videos into three content types (Pro-Recovery, Pro-ED, and ED Experiences) and find that each exhibits distinct emotional interaction patterns. These findings suggest that dense, emotionally structured relationships can emerge within ED discourse on TikTok. More broadly, our work provides one of the first empirical demonstrations of how community-like relational dynamics form and persist on algorithmically driven short-form video platforms.

cs.SI↗

BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok

Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). This content can evade moderation by camouflaging as legitimate fitness advice and disproportionately affects adolescent males. We address this challenge with BigTokDetect, a clinically informed framework for identifying pro-bigorexia content on TikTok. We introduce BigTok, the first expert-annotated multimodal benchmark dataset of over 2,200 TikTok videos labeled by clinical psychiatrists across five categories and eighteen fine-grained subcategories. Comprehensive evaluation of state-of-the-art vision-language models reveals that while commercial zero-shot models achieve the highest accuracy on broad primary categories, supervised fine-tuning enables smaller open-source models to perform better on fine-grained subcategory detection. Ablation studies show that multimodal fusion improves performance by 5 to 15 percent, with video features providing the most discriminative signals. These findings support a grounded moderation approach that automates detection of explicit harms while flagging ambiguous content for human review, and they establish a scalable framework for harm mitigation in emerging mental health domains.

cs.CV↗

Illusions of Intimacy: How Emotional Dynamics Shape Human-AI Relationships

AI companion chatbots, such as those offered by Replika and CharacterAI, increasingly function as always-available companions that provide empathy, validation, and support. While these systems appear to meet basic needs for connection, mounting safety concerns raise a deeper question: how do processes of emotional bonding and intimacy formation unfold in human-AI relationships? Prior research has relied largely on self-reports, interviews, or clinical assessments, leaving unclear how real-world emotional dynamics develop within ongoing human-AI conversations. We address this gap by analyzing over 17,000 user-shared chats with social chatbots from Reddit forums. We show that AI companions dynamically track and mimic user affect and amplify positive emotions, including when users share explicit or transgressive content. These dynamics suggest how chatbots can engage psychological processes involved in intimacy formation and emotional bonding. Finally, we release an anonymized dataset of emotionally salient human-AI companion dialogues to support future empirical work and discuss implications for redesigning and governing social chatbots as high-risk systems for vulnerable users.

cs.SI↗

Leveraging Machine Learning to Identify Gendered Stereotypes and Body Image Concerns on Diet and Fitness Online Forums

The pervasive expectations about ideal body types in Western society can lead to body image concerns, dissatisfaction, and in extreme cases, eating disorders and other psychopathologies related to body image. While previous research has focused on online pro-anorexia communities glorifying the "thin ideal," less attention has been given to the broader spectrum of body image concerns or how emerging disorders like muscle dysmorphia ("bigorexia") present on online platforms. To address this gap, we analyze 46 Reddit forums related to diet, fitness, and mental health. We map these communities along gender and body ideal dimensions, revealing distinct patterns of emotional expression and community support. Feminine-oriented communities, especially those endorsing the thin ideal, express higher levels of negative emotions and receive caring comments in response. In contrast, muscular ideal communities display less negativity, regardless of gender orientation, but receive aggressive compliments in response, marked by admiration and toxicity. Mental health discussions align more with thin ideal, feminine-leaning spaces. By uncovering these gendered emotional dynamics, our findings can inform the development of moderation strategies that foster supportive interactions while reducing exposure to harmful content.

cs.SI↗

EDTok: A Dataset for Eating Disorder Content on TikTok

Eating disorders, which include anorexia nervosa and bulimia nervosa, have been exacerbated by the COVID-19 pandemic, with increased diagnoses linked to heightened exposure to idealized body images online. TikTok, a platform with over a billion predominantly adolescent users, has become a key space where eating disorder content is shared, raising concerns about its impact on vulnerable populations. In response, we present a curated dataset of 43,040 TikTok videos, collected using keywords and hashtags related to eating disorders. Spanning from January 2019 to June 2024, this dataset, offers a comprehensive view of eating disorder-related content on TikTok. Our dataset has the potential to address significant research gaps, enabling analysis of content spread and moderation, user engagement, and the pandemic's influence on eating disorder trends. This work aims to inform strategies for mitigating risks associated with harmful content, contributing valuable insights to the study of digital health and social media's role in shaping mental health.

cs.SI↗

Improving and Assessing the Fidelity of Large Language Models Alignment to Online Communities

Large language models (LLMs) have shown promise in representing individuals and communities, offering new ways to study complex social dynamics. However, effectively aligning LLMs with specific human groups and systematically assessing the fidelity of the alignment remains a challenge. This paper presents a robust framework for aligning LLMs with online communities via instruction-tuning and comprehensively evaluating alignment across various aspects of language, including authenticity, emotional tone, toxicity, and harm. We demonstrate the utility of our approach by applying it to online communities centered on dieting and body image. We administer an eating disorder psychometric test to the aligned LLMs to reveal unhealthy beliefs and successfully differentiate communities with varying levels of eating disorder risk. Our results highlight the potential of LLMs in automated moderation and broader applications in public health and social science research.

cs.CL↗

Safe Spaces or Toxic Places? Content Moderation and Social Dynamics of Online Eating Disorder Communities

Social media platforms have become critical spaces for discussing mental health concerns, including eating disorders. While these platforms can provide valuable support networks, they may also amplify harmful content that glorifies disordered cognition and self-destructive behaviors. While social media platforms have implemented various content moderation strategies, from stringent to laissez-faire approaches, we lack a comprehensive understanding of how these different moderation practices interact with user engagement in online communities around these sensitive mental health topics. This study addresses this knowledge gap through a comparative analysis of eating disorder discussions across Twitter/X, Reddit, and TikTok. Our findings reveal that while users across all platforms engage similarly in expressing concerns and seeking support, platforms with weaker moderation (like Twitter/X) enable the formation of toxic echo chambers that amplify pro-anorexia rhetoric. These results demonstrate how moderation strategies significantly influence the development and impact of online communities, particularly in contexts involving mental health and self-harm.

cs.SI↗

COMMUNITY-CROSS-INSTRUCT: Unsupervised Instruction Generation for Aligning Large Language Models to Online Communities

Social scientists use surveys to probe the opinions and beliefs of populations, but these methods are slow, costly, and prone to biases. Recent advances in large language models (LLMs) enable the creating of computational representations or "digital twins" of populations that generate human-like responses mimicking the population's language, styles, and attitudes. We introduce Community-Cross-Instruct, an unsupervised framework for aligning LLMs to online communities to elicit their beliefs. Given a corpus of a community's online discussions, Community-Cross-Instruct automatically generates instruction-output pairs by an advanced LLM to (1) finetune a foundational LLM to faithfully represent that community, and (2) evaluate the alignment of the finetuned model to the community. We demonstrate the method's utility in accurately representing political and diet communities on Reddit. Unlike prior methods requiring human-authored instructions, Community-Cross-Instruct generates instructions in a fully unsupervised manner, enhancing scalability and generalization across domains. This work enables cost-effective and automated surveying of diverse online communities.

cs.CL↗

Large Language Models Help Reveal Unhealthy Diet and Body Concerns in Online Eating Disorders Communities

Eating disorders (ED), a severe mental health condition with high rates of mortality and morbidity, affect millions of people globally, especially adolescents. The proliferation of online communities that promote and normalize ED has been linked to this public health crisis. However, identifying harmful communities is challenging due to the use of coded language and other obfuscations. To address this challenge, we propose a novel framework to surface implicit attitudes of online communities by adapting large language models (LLMs) to the language of the community. We describe an alignment method and evaluate results along multiple dimensions of semantics and affect. We then use the community-aligned LLM to respond to psychometric questionnaires designed to identify ED in individuals. We demonstrate that LLMs can effectively adopt community-specific perspectives and reveal significant variations in eating disorder risks in different online communities. These findings highlight the utility of LLMs to reveal implicit attitudes and collective mindsets of communities, offering new tools for mitigating harmful content on social media.

cs.SI↗