Searcharxiv⌕ Search

arXiv subjects

Kazutoshi Sasahara

Publications and source records attributed to Kazutoshi Sasahara.

At least 19 recordsLinked to original sources

Uncovering Non-Normality in Information Flow: Network Structure and Dynamics of Social Media Cascades

Information cascades on social media are conventionally conceptualized as directed, feedforward branching processes. However, real-world diffusion pathways frequently deviate from pure hierarchical trees due to localized clustering, reciprocal commentary, and multi-wave temporal surges. In this work, we quantify the directional asymmetry and hierarchical structure of empirical information cascades on X (formerly Twitter) using spectral non-normality via Henrici's departure from normality. Analyzing approximately 58,000 cascade networks across diverse topics (including politics, entertainment, natural disasters, etc.), we investigate (1) how non-normality relates to temporal dynamics such as endogenous-like versus exogenous-like patterns and burstiness, (2) whether non-normality is correlated with the peak concentration or overall size of a cascade, (3) whether the overall non-normality of a cascade's network structure can be predicted from its early stages. We find that non-normality strongly aligns with peak concentration (peak/N) rather than overall cascade size, characterizing cascades governed by rapid, asymmetric forwarding. Furthermore, while early-stage structural forecasting (<= 30% of nodes observed) exhibits expected baseline uncertainty (51%-72% accuracy at a +/- 20% error tolerance), predictability consolidates rapidly during intermediate growth, exceeding 80% across all dynamic clusters once 50%-60% of the network is observed. By identifying the topological and dynamic correlates of cascade structures, this study advances our understanding of information flow and establishes a quantifiable benchmark for forecasting directional diffusion architectures.

cs.SI↗

A latent dimension of Condorcet's jury theorem for multiple AI advisers

When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the majority more reliable. The theorem, however, has a latent dimension when viewed from the user's vantage: adding advisers also makes disagreement more visible. A binomial model reveals that this ``visible dissent'' becomes nearly inevitable as the number of advisers grows, and that reliability and disagreement approach certainty at rates that cross at an adviser accuracy of 4/5 (0.8); below it, visible dissent eventually becomes more likely than a correct majority. Even ideal panels of independent and competent advisers can be correct in aggregate but appear divided; such disagreement does not by itself indicate aggregation failure. The way advisers split also provides a common basis for predictive multiplicity, reconciliation load, and reliance miscalibration. These results separate aggregation from disclosure and turn the latter into testable questions about how disagreement should be presented and interpreted.

cs.CY↗

Attentional DoS: Repeat Reposting, Collective Attention, and Information Diffusion on X

Collective attention is a finite shared resource, and social media posts compete for limited opportunities to be seen. On X, users can undo a repost and repost it again. By repeating this cycle, a user can put the same post back into followers' timelines any number of times without making new content. We call this procedure repeat reposting, and we read it as placing repeated demand on this shared resource (Attentional DoS). We formalize this idea and explore repeat reposting in a large-scale dataset of cascades with at least 1,000 reactions, originating from posts classified as Japanese on X, covering April 2025 to March 2026. Repeat reposts are rare, appearing in a small share of all (user, post) pairs. Even so, close to a million posts have at least one repeater, and most repeats come from a small group of habitual accounts. The central result is that subsequent audience growth is associated less with the number of repeats than with the estimated reach of the repeating accounts. Through the lens of seriality, how habitually the same groups of accounts repeat reposts across many posts, a distinct distributed form emerges: several serial amplifiers converge on the same post (Attentional DDoS). Its synchrony, how closely their actions are timed together, forms a continuum from bursts on the scale of minutes to a daily clock. A check of the content of 99.7% of amplified posts shows that most ADDoS repeat events are directed at Chinese-script content, which our vocabulary matching and sample inspection indicate is predominantly commercial spam. Repeat reposting thus lets a post re-enter the competition for visibility. The observed pattern is better characterized as repeated temporal coverage of an existing audience and the self-reinforcement of posts that are already growing, rather than as evidence that repeat reposting takes reach from other content.

cs.SI↗

Crude, Commercial, and Self-Referential: Chinese-Language Coordinated Activity in Japanese-Language X

Malicious coordination has long been regarded as a principal source of information ecosystem pollution. Here, we focus on crude, text-repetition-based coordination. As the demand for mitigating its dissemination has grown, scholars have studied such coordination, focusing especially on bot detection. Few studies, however, have characterized malicious coordination per se or examined how it elicits reactions from general users. Leveraging a dataset of 734,173 Chinese-language coordinated accounts and around 495 million coordinated posts published between May 2024 and March 2026, this study analyzes the characteristics of coordinated behavior and how general users react to coordinated posts. We report three findings: (1) most coordinated accounts are crude and retain the classic marks of automation, and the same criterion applied to Japanese-language accounts over the same month yields a share six times lower; (2) their content is overwhelmingly non-political; (3) regarding their reach, most reactions within large observable cascades originate from coordinated accounts themselves, while posts classified as potentially harmful or illegal material receive a comparatively high proportion of reactions from outside the Chinese-dominant population. We provide a longitudinal quantitative map of crude Chinese-language coordination appearing in X's Japanese-classified stream.

cs.SI↗

Segregation Before Polarization: How Recommendation Strategies Shape Echo Chamber Pathways

Social media platforms facilitate echo chambers through feedback loops between user preferences and recommendation algorithms. While algorithmic homogeneity is well-documented, the distinct evolutionary pathways driven by content-based versus link-based recommendations remain unclear. Using an extended dynamic Bounded Confidence Model (BCM), we show that content-based algorithms -- unlike their link-based counterparts -- steer social networks toward a segregation-before-polarization (SbP) pathway. Along this trajectory, structural segregation precedes opinion divergence, accelerating individual isolation while delaying but ultimately intensifying collective polarization. Furthermore, we reveal that reposting appears connective by circulating content beyond direct follow links, yet it simultaneously reinforces echo chambers because it amplifies small, latent opinion differences that would otherwise remain inconsequential. These findings suggest that mitigating polarization requires stage-dependent algorithmic interventions, shifting from content-centric to structure-centric strategies as networks evolve.

cs.SI↗

Domain-based user embedding for competing events on social media

Social divide and polarization have become significant societal issues. To understand the mechanisms behind these phenomena, social media analysis offers research opportunities in computational social science, where developing effective user embedding methods is essential for subsequent analysis. Traditionally, researchers have used predefined network-based user features (e.g., network size, degree, and centrality measures). However, because such measures may not capture the complex characteristics of social media users, in our study we developed a method for embedding users based on a URL domain co-occurrence network. This approach effectively represents social media users involved in competing events such as political campaigns and public health crises. We assessed the method's performance using binary classification tasks and datasets that covered topics associated with the COVID-19 infodemic, such as QAnon, Biden, and Ivermectin, among Twitter users. Our results revealed that user embeddings generated directly from the retweet network and/or based on language performed below expectations, whereas our domain-based embeddings outperformed those methods while reducing computation time. Therefore, domain-based embedding offers an accessible and effective method for characterizing social media users in competing events.

cs.CY↗

A multidisciplinary framework for deconstructing bots' pluripotency in dualistic antagonism

Anthropomorphic social bots are engineered to emulate human verbal communication and generate toxic or inflammatory content across social networking services (SNSs). Bot-disseminated misinformation could subtly yet profoundly reshape societal processes by complexly interweaving factors like repeated disinformation exposure, amplified political polarization, compromised indicators of democratic health, shifted perceptions of national identity, propagation of false social norms, and manipulation of collective memory over time. However, extrapolating bots' pluripotency across hybridized, multilingual, and heterogeneous media ecologies from isolated SNS analyses remains largely unknown, underscoring the need for a comprehensive framework to characterise bots' emergent risks to civic discourse. Here we propose an interdisciplinary framework to characterise bots' pluripotency, incorporating quantification of influence, network dynamics monitoring, and interlingual feature analysis. When applied to the geopolitical discourse around the Russo-Ukrainian conflict, results from interlanguage toxicity profiling and network analysis elucidated spatiotemporal trajectories of pro-Russian and pro-Ukrainian human and bots across hybrid SNSs. Weaponized bots predominantly inhabited X, while human primarily populated Reddit in the social media warfare. This rigorous framework promises to elucidate interlingual homogeneity and heterogeneity in bots' pluripotent behaviours, revealing synergistic human-bot mechanisms underlying regimes of information manipulation, echo chamber formation, and collective memory manifestation in algorithmically structured societies.

cs.CY↗

Detecting directional forces in the evolution of grammar: A case study of the English perfect with intransitives across EEBO, COHA, and Google Books

Languages have diverse characteristics that have emerged through evolution. In modern English grammar, the perfect is formed with \textit{have}+PP (past participle), but in earlier English the \textit{be}+PP form also existed. It is widely recognised that the auxiliary verb BE was replaced by HAVE throughout evolution, except for some special cases. However, whether this evolution was caused by natural selection or random drift is still unclear. Here we examined directional forces in the evolution of the English perfect with intransitives by combining three large-scale data sources: EEBO (Early English Books Online), COHA (Corpus of Historical American English), and Google Books. We found that most intransitive verbs exhibited an apparent transition from \textit{be}+PP to \textit{have}+PP, most of which were classified as `selection' by a deep neural network-based model. These results suggest that the English perfect could have evolved through natural selection rather than random drift, and provide insights into the cultural evolution of grammar.

cs.CY↗

Applicability of Trust Management Algorithm in C2C services

The emergence of Consumer-to-Consumer (C2C) platforms has allowed consumers to buy and sell goods directly, but it has also created problems, such as commodity fraud and fake reviews. Trust Management Algorithms (TMAs) are expected to be a countermeasure to detect fraudulent users. However, it is unknown whether TMAs are as effective as reported as they are designed for Peer-to-Peer (P2P) communications between devices on a network. Here we examine the applicability of `EigenTrust', a representative TMA, for the use case of C2C services using an agent-based model. First, we defined the transaction process in C2C services, assumed six types of fraudulent transactions, and then analysed the dynamics of EigenTrust in C2C systems through simulations. We found that EigenTrust could correctly estimate low trust scores for two types of simple frauds. Furthermore, we found the oscillation of trust scores for two types of advanced frauds, which previous research did not address. This suggests that by detecting such oscillations, EigenTrust may be able to detect some (but not all) advanced frauds. Our study helps increase the trustworthiness of transactions in C2C services and provides insights into further technological development for consumer services.

cs.CR↗

Moral intuitions behind deepfake-related discussions in Reddit communities

Deepfakes are AI-synthesized content that are becoming popular on many social media platforms, meaning the use of deepfakes is increasing in society, regardless of its societal implications. Its implications are harmful if the moral intuitions behind deepfakes are problematic; thus, it is important to explore how the moral intuitions behind deepfakes unfold in communities at scale. However, understanding perceived moral viewpoints unfolding in digital contexts is challenging, due to the complexities in conversations. In this research, we demonstrate how Moral Foundations Theory (MFT) can be used as a lens through which to operationalize moral viewpoints in discussions about deepfakes on Reddit communities. Using the extended Moral Foundations Dictionary (eMFD), we measured the strengths of moral intuition (moral loading) behind 101,869 Reddit posts. We present the discussions that unfolded on Reddit in 2018 to 2022 wherein intuitions behind some posts were found to be morally questionable to society. Our results may help platforms detect and take action against immoral activities related to deepfakes.

cs.CY↗

"This is Fake News": Characterizing the Spontaneous Debunking from Twitter Users to COVID-19 False Information

False information spreads on social media, and fact-checking is a potential countermeasure. However, there is a severe shortage of fact-checkers; an efficient way to scale fact-checking is desperately needed, especially in pandemics like COVID-19. In this study, we focus on spontaneous debunking by social media users, which has been missed in existing research despite its indicated usefulness for fact-checking and countering false information. Specifically, we characterize the tweets with false information, or fake tweets, that tend to be debunked and Twitter users who often debunk fake tweets. For this analysis, we create a comprehensive dataset of responses to fake tweets, annotate a subset of them, and build a classification model for detecting debunking behaviors. We find that most fake tweets are left undebunked, spontaneous debunking is slower than other forms of responses, and spontaneous debunking exhibits partisanship in political topics. These results provide actionable insights into utilizing spontaneous debunking to scale conventional fact-checking, thereby supplementing existing research from a new perspective.

cs.SI↗

A network-based approach to QAnon user dynamics and topic diversity during the COVID-19 infodemic

QAnon is an umbrella conspiracy theory that encompasses a wide spectrum of people. The COVID-19 pandemic has helped raise the QAnon conspiracy theory to a wide-spreading movement, especially in the US. Here, we study users' dynamics on Twitter related to the QAnon movement (i.e., pro-/anti-QAnon and less-leaning users) in the context of the COVID-19 infodemic and the topics involved using a simple network-based approach. We found that pro- and anti-leaning users show different population dynamics and that late less-leaning users were mostly anti-QAnon. These trends might have been affected by Twitter's suspension strategies. We also found that QAnon clusters include many bot users. Furthermore, our results suggest that QAnon continues to evolve amid the infodemic and does not limit itself to its original idea but instead extends its reach to create a much larger umbrella conspiracy theory. The network-based approach in this study is important for nowcasting the evolution of the QAnon movement.

cs.SI↗

Psycho-linguistic differences among competing vaccination communities on social media

Currently, the significance of social media in disseminating noteworthy information on topics such as health, politics, and the economy is indisputable. During the COVID-19 pandemic, anti-vaxxers use social media to distribute fake news and anxiety-provoking information about the vaccine, which may harm the public. Here, we characterize the psycho-linguistic features of anti-vaxxers on the online social network Twitter. For this, we collected COVID-19 related tweets from February 2020 to June 2021 to analyse vaccination stance, linguistic features, and social network characteristics. Our results demonstrated that, compared to pro-vaxxers, anti-vaxxers tend to have more negative emotions, narrative thinking, and worse moral tendencies. This study can advance our understanding of the online anti-vaccination movement, and become critical for social media management and policy action during and after the pandemic.

cs.CY↗

Are Deepfakes Concerning? Analyzing Conversations of Deepfakes on Reddit and Exploring Societal Implications

Deepfakes are synthetic content generated using advanced deep learning and AI technologies. The advancement of technology has created opportunities for anyone to create and share deepfakes much easier. This may lead to societal concerns based on how communities engage with it. However, there is limited research available to understand how communities perceive deepfakes. We examined deepfake conversations on Reddit from 2018 to 2021 -- including major topics and their temporal changes as well as implications of these conversations. Using a mixed-method approach -- topic modeling and qualitative coding, we found 6,638 posts and 86,425 comments discussing concerns of the believable nature of deepfakes and how platforms moderate them. We also found Reddit conversations to be pro-deepfake and building a community that supports creating and sharing deepfake artifacts and building a marketplace regardless of the consequences. Possible implications derived from qualitative codes indicate that deepfake conversations raise societal concerns. We propose that there are implications for Human Computer Interaction (HCI) to mitigate the harm created from deepfakes.

cs.HC↗

Characterizing the Anti-Vaxxers' Reply Behavior on Social Media

Although the online campaigns of anti-vaccine advocates, or anti-vaxxers, severely threaten efforts for herd immunity, their reply behavior--the form of directed messaging that can be sent beyond follow-follower relationships--remains poorly understood. Here, we examined the characteristics of anti-vaxxers' reply behavior on Twitter to attempt to comprehend their characteristics of spreading their beliefs in terms of interaction frequency, content, and targets. Among the results, anti-vaxxers more frequently conducted reply behavior with other clusters, especially neutral accounts. Anti-vaxxers' replies were significantly more toxic than those from neutral accounts and pro-vaxxers, and their toxicity, in particular, was higher with regard to the rollout of vaccines. Anti-vaxxers' replies were more persuasive than the others in terms of the emotional aspect, rather than linguistical styles. The targets of anti-vaxxers' replies tend to be accounts with larger numbers of followers and posts, including accounts that relate to health care or represent scientists, policy-makers, or media figures or outlets. We discussed how their reply behaviors are effective in spreading their beliefs, as well as possible countermeasures to restrain them. These findings should prove useful for pro-vaxxers and platformers to promote trusted information while reducing the effect of vaccine disinformation.

cs.SI↗

Morality-based Assertion and Homophily on Social Media: A Cultural Comparison between English and Japanese Languages

Moral psychology is a domain that deals with moral identity, appraisals and emotions. Previous work has primarily focused on moral development and the associated role of culture. Knowing that language is an inherent element of a culture, we used the social media platform Twitter to compare moral behaviors of Japanese tweets with English tweets. The five basic moral foundations, i.e., Care, Fairness, Ingroup, Authority and Purity, along with the associated emotional valence were compared between English and Japanese tweets. The tweets from Japanese users depicted relatively higher Fairness, Ingroup, and Purity, whereas English tweets expressed more positive emotions for all moral dimensions. Considering moral similarities in connecting users on social media, we quantified homophily concerning different moral dimensions using our proposed method. The moral dimensions Care, Authority and Purity for English and Ingroup, Authority and Purity for Japanese depicted homophily on Twitter. Overall, our study uncovers the underlying cultural differences with respect to moral behavior in English- and Japanese-speaking users.

cs.CL↗

Characterizing the roles of bots during the COVID-19 infodemic on Twitter

An infodemic is an emerging phenomenon caused by an overabundance of information online. This proliferation of information makes it difficult for the public to distinguish trustworthy news and credible information from untrustworthy sites and non-credible sources. The perils of an infodemic debuted with the outbreak of the COVID-19 pandemic and bots (i.e., automated accounts controlled by a set of algorithms) that are suspected of spreading the infodemic. Although previous research has revealed that bots played a central role in spreading misinformation during major political events, how bots behaved during the infodemic is unclear. In this paper, we examined the roles of bots in the case of the COVID-19 infodemic and the diffusion of non-credible information such as "5G" and "Bill Gates" conspiracy theories and content related to "Trump" and "WHO" by analyzing retweet networks and retweeted items. We show the segregated topology of their retweet networks, which indicates that right-wing self-media accounts and conspiracy theorists may lead to this opinion cleavage, while malicious bots might favor amplification of the diffusion of non-credible information. Although the basic influence of information diffusion could be larger in human users than bots, the effects of bots are non-negligible under an infodemic situation.

cs.CY↗

Social Influence and Unfollowing Accelerate the Emergence of Echo Chambers

While social media make it easy to connect with and access information from anyone, they also facilitate basic influence and unfriending mechanisms that may lead to segregated and polarized clusters known as "echo chambers." Here we study the conditions in which such echo chambers emerge by introducing a simple model of information sharing in online social networks with the two ingredients of influence and unfriending. Users can change both their opinions and social connections based on the information to which they are exposed through sharing. The model dynamics show that even with minimal amounts of influence and unfriending, the social network rapidly devolves into segregated, homogeneous communities. These predictions are consistent with empirical data from Twitter. Although our findings suggest that echo chambers are somewhat inevitable given the mechanisms at play in online social media, they also provide insights into possible mitigation strategies.

cs.CY↗