SearcharxivSearch

arXiv subjects

Shayan Alipour

Publications and source records attributed to Shayan Alipour.

6 recordsLinked to original sources

When Attention Becomes Exposure in Generative Search

Generative search engines are reshaping information access by replacing traditional ranked lists with synthesized answers and references. In parallel, with the growth of Web3 platforms, incentive-driven creator ecosystems have become an essential part of how enterprises build visibility and community by rewarding creators for contributing to shared narratives. However, the extent to which exposure in generative search engine citations is shaped by external attention markets remains uncertain. In this study, we audit the exposure for 44 Web3 enterprises. First, we show that the creator community around each enterprise is persistent over time. Second, enterprise-specific queries reveal that more popular voices systematically receive greater citation exposure than others. Third, we find that larger follower bases and enterprises with more concentrated creator cores are associated with higher-ranked exposure. Together, these results show that generative search engine citations exhibit exposure bias toward already prominent voices, which risks entrenching incumbents and narrowing viewpoint diversity.

cs.IR

The Gray Area: Characterizing Moderator Disagreement on Reddit

Volunteer moderators play a crucial role in sustaining online dialogue, but they often disagree about what should or should not be allowed. In this paper, we study the complexity of content moderation with a focus on disagreements between moderators, which we term the ``gray area'' of moderation. Leveraging 5 years and 4.3 million moderation log entries from 24 subreddits of different topics and sizes, we characterize how gray area, or disputed cases, differ from undisputed cases. We show that one-in-seven moderation cases are disputed among moderators, often addressing transgressions where users' intent is not directly legible, such as in trolling and brigading, as well as tensions around community governance. This is concerning, as almost half of all gray area cases involved automated moderation decisions. Through information-theoretic evaluations, we demonstrate that gray area cases are inherently harder to adjudicate than undisputed cases and show that state-of-the-art language models struggle to adjudicate them. We highlight the key role of expert human moderators in overseeing the moderation process and provide insights about the challenges of current moderation processes and tools.

cs.CY

Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness

Large language models (LLMs) are known to exhibit demographic biases, yet few studies systematically evaluate these biases across multiple datasets or account for confounding factors. In this work, we examine LLM alignment with human annotations in five offensive language datasets, comprising approximately 220K annotations. Our findings reveal that while demographic traits, particularly race, influence alignment, these effects are inconsistent across datasets and often entangled with other factors. Confounders -- such as document difficulty, annotator sensitivity, and within-group agreement -- account for more variation in alignment patterns than demographic traits alone. Specifically, alignment increases with higher annotator sensitivity and group agreement, while greater document difficulty corresponds to reduced alignment. Our results underscore the importance of multi-dataset analyses and confounder-aware methodologies in developing robust measures of demographic bias in LLMs.

cs.CY

Cross-Platform Social Dynamics: An Analysis of ChatGPT and COVID-19 Vaccine Conversations

The role of social media in information dissemination and agenda-setting has significantly expanded in recent years. By offering real-time interactions, online platforms have become invaluable tools for studying societal responses to significant events as they unfold. However, online reactions to external developments are influenced by various factors, including the nature of the event and the online environment. This study examines the dynamics of public discourse on digital platforms to shed light on this issue. We analyzed over 12 million posts and news articles related to two significant events: the release of ChatGPT in 2022 and the global discussions about COVID-19 vaccines in 2021. Data was collected from multiple platforms, including Twitter, Facebook, Instagram, Reddit, YouTube, and GDELT. We employed topic modeling techniques to uncover the distinct thematic emphases on each platform, which reflect their specific features and target audiences. Additionally, sentiment analysis revealed various public perceptions regarding the topics studied. Lastly, we compared the evolution of engagement across platforms, unveiling unique patterns for the same topic. Notably, discussions about COVID-19 vaccines spread more rapidly due to the immediacy of the subject, while discussions about ChatGPT, despite its technological importance, propagated more gradually.

cs.CY

The Drivers of Global News Spreading Patterns

The web radically changed the dissemination of information and the global spread of news. In this study, we aim to reconstruct the connectivity patterns within nations shaping news propagation globally in 2022. We do this by analyzing a dataset of unprecedented size, containing 140 million news articles from 183 countries and related to 37,802 domains in the GDELT database. Unlike previous research, we focus on the sequential mention of events across various countries, thus incorporating a temporal dimension into the analysis of news dissemination networks. Our results show a significant imbalance in online news spreading. We identify news superspreaders forming a tightly interconnected rich club, exerting significant influence on the global news agenda. To further investigate the mechanisms underlying news dissemination and the shaping of global public opinion, we model countries' interactions using a gravity model, incorporating economic, geographical, and cultural factors. Consistent with previous studies, we find that countries' GDP is one of the main drivers to shape the worldwide news agenda.

cs.SI

Users volatility on Reddit and Voat

Social media platforms are like giant arenas where users can rely on different content and express their opinions through likes, comments, and shares. However, do users welcome different perspectives or only listen to their preferred narratives? This paper examines how users explore the digital space and allocate their attention among communities on two social networks, Voat and Reddit. By analysing a massive dataset of about 215 million comments posted by about 16 million users on Voat and Reddit in 2019 we find that most users tend to explore new communities at a decreasing rate, meaning they have a limited set of preferred groups they visit regularly. Moreover, we provide evidence that preferred communities of users tend to cover similar topics throughout the year. We also find that communities have a high turnover of users, meaning that users come and go frequently showing a high volatility that strongly departs from a null model simulating users' behaviour.

physics.soc-ph