SearcharxivSearch

arXiv subjects

Fujio Toriumi

Publications and source records attributed to Fujio Toriumi.

At least 19 recordsLinked to original sources

From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X

Ragebait refers to online content intentionally designed to provoke anger or outrage and thereby increase attention and engagement. However, reliable large-scale detection and systematic analysis of ragebait remain limited, hindering efforts to understand its prevalence, impact, and mitigation. This study aims to develop an effective ragebait detection framework and to clarify the characteristics of ragebait at scale, providing a basis for understanding and mitigating emotionally provocative content online. We constructed a labeled dataset with the assistance of a large language model (LLM) and trained several Japanese language models for ragebait detection. The resulting ensemble classifier was then applied to a large-scale dataset of Japanese-language posts on X. Our analysis shows that ragebait is more prevalent in politically and socially contentious topics, including politics, discrimination, public health, and interpersonal conflict. Ragebait posts also spread faster and receive more negative reactions than non-ragebait posts, particularly anger, fear, disgust, sadness, and surprise. These findings demonstrate the utility of the proposed detector and provide a large-scale characterization of ragebait in Japanese online discourse.

cs.SI

Network-distance decay of perceived online social support

Perceived social support can buffer stress, but how it is associated across online social networks at different graph distances remains unclear. Here we show that inferred perceived online social support in a large avatar communication application decays with network distance in a form better described, over the observed range, by a power-decay model than by a single exponential. We linked two-wave survey data from Pigg Party with behavioral logs, trained a random forest to infer perceived support for active users, and regressed Wave 2 scores on Wave 1 scores for users at hop distance $k$, adjusting for baseline support and covariates. Adjusted associations persisted across hops, consistent with power-law-like decay over the observed range. Individual-based simulations indicated that heterogeneous source-specific exponential decay rates can generate heavy-tailed aggregate decay. These results suggest that network position heterogeneity should be considered when characterizing distance-dependent associations among psychosocial states in online communities.

cs.SI

Temporal Shifts and Causal Interactions of Emotions in Social and Mass Media: A Case Study of the "Reiwa Rice Riot" in Japan

In Japan, severe rice shortages in 2024 sparked widespread public controversy across both news media and social platforms, culminating in what has been termed the "Reiwa Rice Riot." This study proposes a framework to analyze the temporal dynamics and causal interactions of emotions expressed on X (formerly Twitter) and in news articles, using the "Reiwa Rice Riot" as a case study. While recent studies have shown that emotions mutually influence each other between social and mass media, the patterns and transmission pathways of such emotional shifts remain insufficiently understood. To address this gap, we applied a machine learning-based emotion classification grounded in Plutchik's eight basic emotions to analyze posts from X and domestic news articles. Our findings reveal that emotional shifts and information dissemination on X preceded those in news media. Furthermore, in both media platforms, the fear was initially the most dominant emotion, but over time intersected with hope which ultimately became the prevailing emotion. Our findings suggest that patterns in emotional expressions on social media may serve as a lens for exploring broader social dynamics.

cs.SI

Impact of seed node position on network robustness under localized attacks

Localized attacks (LAs), where damage propagates from a single seed node to its neighbors, pose significant threats to the robustness of complex networks. Although previous studies have extensively analyzed network vulnerability under such attacks, they typically assume random seed node placement and evaluate average robustness. However, the structural position of the seed node can significantly impact the extent of damage. This study proposes the Localized Attack Vulnerability Index (LAVI), a node-level metric that quantifies the potential impact of a LA initiated at a specific node. LAVI quantifies the cumulative number of severed links during attack progression, capturing how local connectivity and topological position amplify the resulting damage. Numerical experiments on synthetic and real-world networks demonstrate that LAVI correlates more strongly with network robustness degradation than standard centrality measures, such as degree, closeness, and betweenness. Our findings highlight that classical centrality metrics fail to capture key dynamics of spatially localized failures, while LAVI provides an accurate and generalizable indicator of node vulnerability under such disruptions.

physics.soc-ph

An attention economy model of co-evolution between content quality and audience selectivity

Human attention has become a scarce and strategically contested resource in digital environments. Content providers increasingly engage in excessive competition for visibility, often prioritizing attention-grabbing tactics over substantive quality. Despite extensive empirical evidence, however, there is a lack of theoretical models that explain the fundamental dynamics of the attention economy. Here, we develop a minimal mathematical framework to explain how content quality and audience attention coevolve under limited attention capacity. Using an evolutionary game approach, we model strategic feedback between providers, who decide how much effort to invest in production, and consumers, who choose whether to search selectively for high-quality content or to engage passively. Analytical and numerical results reveal three characteristic regimes of content dynamics: collapse, boundary, and coexistence. The transitions between these regimes depend on how effectively audiences can distinguish content quality. When audience discriminability is weak, both selective attention and high-quality production vanish, leading to informational collapse. When discriminability is sufficient and incentives are well aligned, high- and low-quality content dynamically coexist through feedback between audience selectivity and providers' effort. These findings identify two key conditions for sustaining a healthy information ecosystem: adequate discriminability among audiences and sufficient incentives for high-effort creation. The model provides a theoretical foundation for understanding how institutional and platform designs can prevent the degradation of content quality in the attention economy.

physics.soc-ph

Global Patterns of Knowledge: Language, Genre, and the Geography of Knowledge

Online platforms, particularly Wikipedia, have become critical infrastructures for providing diverse linguistic and cultural contexts. This human-curated knowledge now forms the foundation for modern AI. However, we have not yet fully explored how knowledge production capability vary across languages and domains. Here, we address this gap by applying economic complexity analysis to understand the editing history of Wikipedia platforms. This approach allows us to infer the latent mode of ``knowledge-production'' of each language community from the diversity and specialization of its contributed content. We reveal that different language communities exhibit distinct specializations, particularly in cultural subjects. Furthermore, we map the global landscape of these production modes, finding that the structure of knowledge production strongly reflects geopolitical boundaries. Our findings suggest that while a common mode of knowledge production exists for standardized topics such as science, it is more diverse for cultural topics or controversial subjects such as conspiracy theories. The association between differences in knowledge production capability and geopolitical factors implies how linguistic and cultural dynamics shape our worldview and the biases embedded in Wikipedia data, a unique, massive, and essential dataset for modern AI.

cs.CY

Reducing Sexual Predation and Victimization Through Warnings and Awareness among High-Risk Users

Online sexual predators target children by building trust, creating dependency, and arranging meetings for sexual purposes. This poses a significant challenge for online communication platforms that strive to monitor and remove such content and terminate predators' accounts. However, these platforms can only take such actions if sexual predators explicitly violate the terms of service, not during the initial stages of relationship-building. This study designed and evaluated a strategy to prevent sexual predation and victimization by delivering warnings and raising awareness among high-risk individuals based on the routine activity theory in criminal psychology. We identified high-risk users as those with a high probability of committing or being subjected to violations, using a machine learning model that analyzed social networks and monitoring data from the platform. We conducted a randomized controlled trial on a Japanese avatar-based communication application, Pigg Party. High-risk players in the intervention group received warnings and awareness-building messages, while those in the control group did not receive the messages, regardless of their risk level. The trial involved 12,842 high-risk players in the intervention group and 12,844 in the control group for 138 days. The intervention successfully reduced violations and being violated among women for 12 weeks, although the impact on men was limited. These findings contribute to efforts to combat online sexual abuse and advance understanding of criminal psychology.

cs.SI

Informational Health --Toward the Reduction of Risks in the Information Space

The modern information society, markedly influenced by the advent of the internet and subsequent developments such as WEB 2.0, has seen an explosive increase in information availability, fundamentally altering human interaction with information spaces. This transformation has facilitated not only unprecedented access to information but has also raised significant challenges, particularly highlighted by the spread of ``fake news'' during critical events like the 2016 U.S. presidential election and the COVID-19 pandemic. The latter event underscored the dangers of an ``infodemic,'' where the large amount of information made distinguishing between factual and non-factual content difficult, thereby complicating public health responses and posing risks to democratic processes. In response to these challenges, this paper introduces the concept of ``informational health,'' drawing an analogy between dietary habits and information consumption. It argues that just as balanced diets are crucial for physical health, well-considered nformation behavior is essential for maintaining a healthy information environment. This paper proposes three strategies for fostering informational health: literacy education, visualization of meta-information, and informational health assessments. These strategies aim to empower users and platforms to navigate and enhance the information ecosystem effectively. By focusing on long-term informational well-being, we highlight the necessity of addressing the social risks inherent in the current attention economy, advocating for a paradigm shift towards a more sustainable information consumption model.

cs.CY

QANA: LLM-based Question Generation and Network Analysis for Zero-shot Key Point Analysis and Beyond

The proliferation of social media has led to information overload and increased interest in opinion mining. We propose "Question-Answering Network Analysis" (QANA), a novel opinion mining framework that utilizes Large Language Models (LLMs) to generate questions from users' comments, constructs a bipartite graph based on the comments' answerability to the questions, and applies centrality measures to examine the importance of opinions. We investigate the impact of question generation styles, LLM selections, and the choice of embedding model on the quality of the constructed QA networks by comparing them with annotated Key Point Analysis datasets. QANA achieves comparable performance to previous state-of-the-art supervised models in a zero-shot manner for Key Point Matching task, also reducing the computational cost from quadratic to linear. For Key Point Generation, questions with high PageRank or degree centrality align well with manually annotated key points. Notably, QANA enables analysts to assess the importance of key points from various aspects according to their selection of centrality measure. QANA's primary contribution lies in its flexibility to extract key points from a wide range of perspectives, which enhances the quality and impartiality of opinion mining.

cs.CL

HyperS2V: A Framework for Structural Representation of Nodes in Hyper Networks

In contrast to regular (simple) networks, hyper networks possess the ability to depict more complex relationships among nodes and store extensive information. Such networks are commonly found in real-world applications, such as in social interactions. Learning embedded representations for nodes involves a process that translates network structures into more simplified spaces, thereby enabling the application of machine learning approaches designed for vector data to be extended to network data. Nevertheless, there remains a need to delve into methods for learning embedded representations that prioritize structural aspects. This research introduces HyperS2V, a node embedding approach that centers on the structural similarity within hyper networks. Initially, we establish the concept of hyper-degrees to capture the structural properties of nodes within hyper networks. Subsequently, a novel function is formulated to measure the structural similarity between different hyper-degree values. Lastly, we generate structural embeddings utilizing a multi-scale random walk framework. Moreover, a series of experiments, both intrinsic and extrinsic, are performed on both toy and real networks. The results underscore the superior performance of HyperS2V in terms of both interpretability and applicability to downstream tasks.

cs.SI

User's Position-Dependent Strategies in Consumer-Generated Media with Monetary Rewards

Numerous forms of consumer-generated media (CGM), such as social networking services (SNS), are widely used. Their success relies on users' voluntary participation, often driven by psychological rewards like recognition and connection from reactions by other users. Furthermore, a few CGM platforms offer monetary rewards to users, serving as incentives for sharing items such as articles, images, and videos. However, users have varying preferences for monetary and psychological rewards, and the impact of monetary rewards on user behaviors and the quality of the content they post remains unclear. Hence, we propose a model that integrates some monetary reward schemes into the SNS-norms game, which is an abstraction of CGM. Subsequently, we investigate the effect of each monetary reward scheme on individual agents (users), particularly in terms of their proactivity in posting items and their quality, depending on agents' positions in a CGM network. Our experimental results suggest that these factors distinctly affect the number of postings and their quality. We believe that our findings will help CGM platformers in designing better monetary reward schemes.

cs.SI

Investigating Quantitative-Qualitative Topical Preference: A Comparative Study of Early and Late Engagers in Japanese ChatGPT Conversations

This study investigates engagement patterns related to OpenAI's ChatGPT on Japanese Twitter, focusing on two distinct user groups - early and late engagers, inspired by the Innovation Theory. Early engagers are defined as individuals who initiated conversations about ChatGPT during its early stages, whereas late engagers are those who began participating at a later date. To examine the nature of the conversations, we employ a dual methodology, encompassing both quantitative and qualitative analyses. The quantitative analysis reveals that early engagers often engage with more forward-looking and speculative topics, emphasizing the technological advancements and potential transformative impact of ChatGPT. Conversely, the late engagers intereact more with contemporary topics, focusing on the optimization of existing AI capabilities and considering their inherent limitations. Through our qualitative analysis, we propose a method to measure the proportion of shared or unique viewpoints within topics across both groups. We found that early engagers generally concentrate on a more limited range of perspectives, whereas late engagers exhibit a wider range of viewpoints. Interestingly, a weak correlation was found between the volume of tweets and the diversity of discussed topics in both groups. These findings underscore the importance of identifying semantic bias, rather than relying solely on the volume of tweets, for understanding differences in communication styles between groups within a given topic. Moreover, our versatile dual methodology holds potential for broader applications, such as studying engagement patterns within different user groups, or in contexts beyond ChatGPT.

cs.SI

Effect of Monetary Reward on Users' Individual Strategies Using Co-Evolutionary Learning

Consumer generated media (CGM), such as social networking services rely on the voluntary activity of users to prosper, garnering the psychological rewards of feeling connected with other people through comments and reviews received online. To attract more users, some CGM have introduced monetary rewards (MR) for posting activity and quality articles and comments. However, the impact of MR on the article posting strategies of users, especially frequency and quality, has not been fully analyzed by previous studies, because they ignored the difference in the standpoint in the CGM networks, such as how many friends/followers they have, although we think that their strategies vary with their standpoints. The purpose of this study is to investigate the impact of MR on individual users by considering the differences in dominant strategies regarding user standpoints. Using the game-theoretic model for CGM, we experimentally show that a variety of realistic dominant strategies are evolved depending on user standpoints in the CGM network, using multiple-world genetic algorithm.

cs.SI

How Many Tweets DoWe Need?: Efficient Mining of Short-Term Polarized Topics on Twitter: A Case Study From Japan

In recent years, social media has been criticized for yielding polarization. Identifying emerging disagreements and growing polarization is important for journalists to create alerts and provide more balanced coverage. While recent studies have shown the existence of polarization on social media, they primarily focused on limited topics such as politics with a large volume of data collected in the long term, especially over months or years. While these findings are helpful, they are too late to create an alert immediately. To address this gap, we develop a domain-agnostic mining method to identify polarized topics on Twitter in a short-term period, namely 12 hours. As a result, we find that daily Japanese news-related topics in early 2022 were polarized by 31.6\% within a 12-hour range. We also analyzed that they tend to construct information diffusion networks with a relatively high average degree, and half of the tweets are created by a relatively small number of people. However, it is very costly and impractical to collect a large volume of tweets daily on many topics and monitor the polarization due to the limitations of the Twitter API. To make it more cost-efficient, we also develop a prediction method using machine learning techniques to estimate the polarization level using randomly collected tweets leveraging the network information. Extensive experiments show a significant saving in collection costs compared to baseline methods. In particular, our approach achieves F-score of 0.85, requiring 4,000 tweets, 4x savings than the baseline. To the best of our knowledge, our work is the first to predict the polarization level of the topics with low-resource tweets. Our findings have profound implications for the news media, allowing journalists to detect and disseminate polarizing information quickly and efficiently.

cs.SI

Do you trust experts on Twitter?: Successful correction of COVID-19-related misinformation

This study focuses on how scientifically-correct information is disseminated through social media, and how misinformation can be corrected. We have identified examples on Twitter where scientific terms that have been misused have been rectified and replaced by scientifically-correct terms through the interaction of users. The results show that the percentage of correct terms ("variant" or "COVID-19 variant") being used instead of the incorrect terms ("strain") on Twitter has already increased since the end of December 2020. This was about a month before the release of an official statement by the Japanese Association for Infectious Diseases regarding the correct terminology, and the use of terms on social media was faster than it was in television. Some Twitter users who quickly started using the correct term were more likely to retweet messages sent by leading influencers on Twitter, rather than messages sent by traditional media or portal sites. However, a few Twitter users continued to use wrong terms even after March 2021, even though the use of the correct terms was widespread. Further analysis of their tweets revealed that they were quoting sources that differed from that of other users. This study empirically verified that self-correction occurs even on Twitter, which is often known as a "hotbed for spreading rumors." The results of this study also suggest that influencers with expertise can influence the direction of public opinion on social media and that the media that users usually cite can also affect the possibility of behavioral changes.

cs.SI

Retrospective Analysis of Controversial Subtopics on COVID-19 in Japan

For efficient political decision-making in an emergency situation, a thorough recognition and understanding of the polarized topics is crucial. The cost of unmitigated polarization would be extremely high for the society; therefore, it is desirable to identify the polarizing issues before they become serious. With this in mind, we conducted a retrospective analysis of the polarized subtopics of COVID-19 to obtain insights for future policymaking. To this end, we first propose a framework to comprehensively search for controversial subtopics. We then retrospectively analyze subtopics on COVID-19 using the proposed framework, with data obtained via Twitter in Japan. The results show that the proposed framework can effectively detect controversial subtopics that reflect current reality. Controversial subtopics tend to be about the government, medical matters, economy, and education; moreover, the controversy score had a low correlation with the traditional indicators--scale and sentiment of the subtopics--which suggests that the controversy score is a potentially important indicator to be obtained. We also discussed the difference between subtopics that became highly controversial and ones that did not despite their large scale.

cs.SI

Corrective Information Does Not Necessarily Curb Social Disruption

The spread of misinformation can cause social confusion. The authenticity of information on a social networking service (SNS) is unknown, and false information can be easily spread. Consequently, many studies have been conducted on methods to control the spread of misinformation on social networking sites. However, few studies have examined the impact of the spread of misinformation and its corrections on society. This study models the impact of the reduction of misinformation and the diffusion of corrective information on social disruption, and it identifies the features of this impact. In this study, we analyzed misinformation regarding the shortage of toilet paper during the 2020 COVID-19 epidemic, its corrections, and the excessive purchasing caused by this information. First, we analyze the amount of misinformation and corrective information spread on SNS, and we create a regression model to estimate the real-world impact of misinformation and its correction. This model is used to analyze the change in real-world impact corresponding to the change in the diffusion of misinformation and corrective information. Our analysis shows that the corrective information was spread to a much greater extent than the misinformation. In addition, our model reveals that the corrective information was what caused the excessive purchasing behavior. As a result of our further analysis, we found that the amount of diffusion of corrective information required to minimize the impact on the real world depends on the amount of the diffusion of misinformation.

cs.SI

Analysis of Political Party Twitter Accounts' Retweeters During Japan's 2017 Election

In modern election campaigns, political parties utilize social media to advertise their policies and candidates and to communicate to the electorate. In Japan's latest general election in 2017, the 48th general election for the Lower House, social media, especially Twitter, was actively used. In this paper, we analyze the users who retweeted tweets of political parties on Twitter during the election. Our aim is to clarify what kinds of users are diffusing (retweeting) tweets of political parties. The results indicate that the characteristics of retweeters of the largest ruling party (Liberal Democratic Party of Japan) and the largest opposition party (The Constitutional Democratic Party of Japan) were similar, even though the retweeters did not overlap each other. We also found that a particular opposition party (Japanese Communist Party) had quite different characteristics from other political parties.

cs.SI