SearcharxivSearch

arXiv subjects

Pedro Ramaciotti

Publications and source records attributed to Pedro Ramaciotti.

7 recordsLinked to original sources

Community Notes undermoderate polarizing content by design creating risks in electoral processes

Community Notes (CNs) of X enables users to collaboratively moderate misleading content. To resolve conflicting moderation, CNs infers a latent ideological dimension and selects notes garnering cross-partisan support. As this system is now deployed worldwide, we evaluate its operation across diverse polarization contexts. We analyze all 1.9 million moderation notes receiving 135 million ratings by March 2025, cross-referencing ideological scaling data on 13 countries. Our results show that the CNs algorithm effectively captures the main polarizing dimensions across countries, surfacing notes that garner cross-partisan support. This also means that, by design, CNs systematically under-moderate polarizing content. We analyze notes relating to four recent elections in the US (2024), the UK (2024), France (2024) and Germany (2025) and demonstrate that they are systematically under-moderated when compared to other notes, posing potential risks to civic discourse and electoral processes.

cs.SI

Political attitudes differ but share a common low-dimensional structure across social media and survey data

Does polarization online reflect the state of polarization in society? We study ideological positions and attitudes on several issues in France, a country with documented issue nonalignment. We compare distributions on X/Twitter with a nationally representative sample, focusing on two key properties: ideological polarization and issue alignment. Despite significant issue-wise divergences, positions of both the X population and the nationally representative sample present a similar bi-dimensional structure along two dominant bundles of aligned issues: a Left-Right divide, and a Global-Local divide. We then study how our results vary when accounting for key structural parameters of the online public sphere: activity, popularity, and visibility. We find that the dimensionality of attitude distributions shrinks as ideological polarization increases when selecting more active users. The divergence between political attitudes on social media and in survey data is greatly mediated by the combination of activity and popularity of social media users: users benefiting from the most exposure are also the most representative of the general public. Together, our results shed light on the structural similarities and differences between political attitudes from social media users and the general public.

cs.SI

Mapping the political landscape from data traces: multidimensional opinions of users, politicians and media outlets on X

Studying political activity on social media often requires defining and measuring political stances of users or content. Relevant examples include the study of opinion polarization, or the study of political diversity in online content diets. While many research designs rely on operationalizations best suited for the US setting, few allow addressing more general political systems, in which users and media outlets might exhibit stances on multiple ideology and issue dimensions, going beyond traditional Liberal-Conservative or Left-Right scales. To advance the study of more general online ecosystems, we present a dataset pertaining to a population of X/Twitter users, parliamentarians, and media outlets embedded in a political space spanned by dimensions measuring attitudes towards immigration, the EU, liberal values, elites and institutions, nationalism and the environment, in addition to left-right and liberal-conservative scales. We include indicators of individual activity and popularity: mean number of posts per day, number of followers, and number of followees. We provide several benchmarks validating the positions of these entities and discuss several applications for this dataset.

cs.SI

Recommender system in X inadvertently profiles ideological positions of users

Studies on recommendations in social media have mainly analyzed the quality of recommended items (e.g., their diversity or biases) and the impact of recommendation policies (e.g., in comparison with purely chronological policies). We use a data donation program, collecting more than 2.5 million friend recommendations made to 682 volunteers on X over a year, to study instead how real-world recommenders learn, represent and process political and social attributes of users inside the so-called black boxes of AI systems. Using publicly available knowledge on the architecture of the recommender, we inferred the positions of recommended users in its embedding space. Leveraging ideology scaling calibrated with political survey data, we analyzed the political position of users in our study (N=26,509 among volunteers and recommended contacts) among several attributes, including age and gender. Our results show that the platform's recommender system produces a spatial ordering of users that is highly correlated with their Left-Right positions (Pearson rho=0.887, p-value < 0.0001), and that cannot be explained by socio-demographic attributes. These results open new possibilities for studying the interaction between human and AI systems. They also raise important questions linked to the legal definition of algorithmic profiling in data privacy regulation by blurring the line between active and passive profiling. We explore new constrained recommendation methods enabled by our results, limiting the political information in the recommender as a potential tool for privacy compliance capable of preserving recommendation relevance.

cs.SI

Linear socio-demographic representations emerge in Large Language Models from indirect cues

We investigate how LLMs encode sociodemographic attributes of human conversational partners inferred from indirect cues such as names and occupations. We show that LLMs develop linear representations of user demographics within activation space, wherein stereotypically associated attributes are encoded along interpretable geometric directions. We first probe residual streams across layers of four open transformer-based LLMs (Magistral 24B, Qwen3 14B, GPT-OSS 20B, OLMo2-1B) prompted with explicit demographic disclosure. We show that the same probes predict demographics from implicit cues: names activate census-aligned gender and race representations, while occupations trigger representations correlated with real-world workforce statistics. These linear representations allow us to explain demographic inferences implicitly formed by LLMs during conversation. We demonstrate that these implicit demographic representations actively shape downstream behavior, such as career recommendations. Our study further highlights that models that pass bias benchmark tests may still harbor and leverage implicit biases, with implications for fairness when applied at scale.

cs.AI

Web Crawler Restrictions, AI Training Datasets \& Political Biases

Large language models rely on web-scraped text for training; concurrently, content creators are increasingly blocking AI crawlers to retain control over their data. We analyze crawler restrictions across the top one million most-visited websites since 2023 and examine their potential downstream effects on training data composition. Our analysis reveals growing restrictions, with blocking patterns varying by website popularity and content type. A quarter of the top thousand websites restrict AI crawlers, decreasing to one-tenth across the broader top million. Content type matters significantly: 34.2% of news outlets disallow OpenAI's GPTBot, rising to 55% for outlets with high factual reporting. Additionally, outlets with neutral political positions impose the strongest restrictions (58%), whereas hyperpartisan websites and those with low factual reporting impose fewer restrictions -only 4.1% of right-leaning outlets block access to OpenAI. Our findings suggest that heterogeneous blocking patterns may skew training datasets toward low-quality or polarized content, potentially affecting the capabilities of models served by prominent AI-as-a-Service providers.

cs.SI

Multidimensional political polarization in online social networks

Political polarization in online social platforms is a rapidly growing phenomenon worldwide. Despite their relevance to modern-day politics, the structure and dynamics of polarized states in digital spaces are still poorly understood. We analyze the community structure of a two-layer, interconnected network of French Twitter users, where one layer contains members of Parliament and the other one regular users. We obtain an optimal representation of the network in a four-dimensional political opinion space by combining network embedding methods and political survey data. We find structurally cohesive groups sharing common political attitudes and relate them to the political party landscape in France. The distribution of opinions of professional politicians is narrower than that of regular users, indicating the presence of more extreme attitudes in the general population. We find that politically extreme communities interact less with other groups as compared to more centrist groups. We apply an empirically tested social influence model to the two-layer network to pinpoint interaction mechanisms that can describe the political polarization seen in data, particularly for centrist groups. Our results shed light on the social behaviors that drive digital platforms towards polarization, and uncover an informative multidimensional space to assess political attitudes online.

physics.soc-ph