SearcharxivSearch

arXiv subjects

Yasin Silva

Publications and source records attributed to Yasin Silva.

3 recordsLinked to original sources

Measuring and Detecting Harmful AI Sycophancy

Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.

cs.AI

SPILLOVER: Measuring Cyberbullying NormPropagation on Social Media

While certain aspects of cyberbullying (CB) such as its factors and prevalence have been studied extensively, relatively little attention has been given to specifically how the aggression transfers from comment to comment. This understanding could have important implications for designing better anti-bullying features. In this paper, we study multiple aspects of the nature of this aggression transference in social media sessions. Using data from 32,754 consecutive comment pairs from 430 Instagram sessions, we find that a preceding CB comment substantially raises the odds of the next comment being CB, an effect confirmed by session fixed-effects controls and driven primarily by cross-user spread. We also find that $\text{CB} \to \text{CB}$ pairs are more textually similar than $\text{NoCB} \to \text{CB}$ pairs across five complementary methods, and that this pattern holds under a matched cross-session baseline that rules out shared vocabulary, session toxicity, and session length as confounds. Moreover, non-aggressive replies grow more negative as preceding CB severity increases, a graded pattern consistent with automatic emotional influence below the threshold of overt aggression. These key findings replicate across three independent datasets (Reddit, Wikipedia Detox, and SOCC), with spillover rates that track platform visibility design. Finally, we show that a single binary feature (whether the prior comment was CB) improves prediction over session-level baselines and over a fine-tuned HateBERT classifier, serving as a real-time moderation signal that targets the spreading chain rather than individual offenders.

cs.SI

From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?

The rapid advancement of large language models (LLMs) has opened new possibilities for AI for good applications. As LLMs increasingly mediate online communication, their potential to foster empathy and constructive dialogue becomes an important frontier for responsible AI research. This work explores whether LLMs can serve not only as moderators that detect harmful content, but as mediators capable of understanding and de-escalating online conflicts. Our framework decomposes mediation into two subtasks: judgment, where an LLM evaluates the fairness and emotional dynamics of a conversation, and steering, where it generates empathetic, de-escalatory messages to guide participants toward resolution. To assess mediation quality, we construct a large Reddit-based dataset and propose a multi-stage evaluation pipeline combining principle-based scoring, user simulation, and human comparison. Experiments show that API-based models outperform open-source counterparts in both reasoning and intervention alignment when doing mediation. Our findings highlight both the promise and limitations of current LLMs as emerging agents for online social mediation.

cs.AI