SearcharxivSearch

arXiv subjects

Mitsuo Yoshida

Publications and source records attributed to Mitsuo Yoshida.

At least 19 recordsLinked to original sources

Temporal Portability of Numeric User Metadata on Twitter

Numeric user metadata in social media are often reused over time. However, their reusability may depend on what an analysis needs to preserve. We introduce temporal portability as an analytical perspective for assessing the cross-time reuse of user features and feature-based rules. Specifically, we ask how well relevant properties are preserved when features and rules defined at a source time point are reused at a target time point. We used quarterly data on user features obtained directly from or derived from Japanese-language tweets in Twitter's 1% sample stream from 2020-Q1 to 2022-Q3. Each quarter included approximately 10.1--11.0 million unique users. We evaluated 13 numeric user features in terms of feature distributions, same-user relative ranks, selection rates, and selected-user membership. Across quarters, feature distributions changed and, for many features, same-user relative ranks were less well preserved at longer quarter lags. Reusing source-quarter thresholds also produced selection-rate drift. Target-quarter recalibration nearly matched source-quarter selection rates. However, membership turnover persisted and increased at longer quarter lags. Our results show that temporal portability should be assessed in terms of the property that an analysis needs to preserve.

cs.SI

Assessing Post-Reform Changes in Risk Disclosure Quality with a Multidimensional Text Analysis Approach

While corporate narrative disclosures provide crucial information to capital markets, comprehensively evaluating their qualitative changes over time remains challenging. Narrative text is inherently multidimensional, meaning that an improvement in one textual dimension often occurs alongside changes in others. To capture these underlying dynamics, we propose a longitudinal text analysis approach combining Japanese-language NLP metric extraction with paired testing, shift function analysis, and inter-metric correlation. Our framework extends prior indicator sets by incorporating a cross-section relevance indicator to measure topical alignment between risk disclosures and management strategies. Applying this approach to evaluate Japan's 2019 disclosure reforms, we analyze 19,770 firm-year observations over a 10-year period (FY2015-FY2024). The joint analysis reveals complex shifts in disclosure patterns that are frequently masked by conventional single-indicator methods. Specifically, we find that while disclosure volume increased substantially, it was accompanied by a decline in readability. Furthermore, although the overall information structure improved, specific descriptive quality stagnated, and the degree of adaptation varied across market segments.

cs.CL

Mapping Social Media User Behaviors in Reciprocity Space

Social media users exhibit diverse behavioral patterns as platforms function simultaneously as information and friendship networks. We introduce a reciprocity-based framework mapping users onto two-dimensional space defined by bidirectional connection ratios. Analyzing 48,830 Twitter users and 149 million connections, we demonstrate that fragmented user types from prior studies (influencers, lurkers, brokers, and follow-back accounts) emerge naturally as regions within continuous behavioral space rather than discrete categories. User properties vary smoothly across the reciprocity dimensions, revealing clear behavioral gradients. This framework provides the first unified model encompassing the full spectrum of social media behaviors and offers interpretable metrics for influence measurement and platform design.

cs.SI

Identifying Stable Influencers: Distinguishing Stable and Temporal Influencers Using Long-Term Twitter Data

For effective social media marketing, identifying stable influencers-those who sustain their influence over an extended period-is more valuable than focusing on users who are influential only temporarily. This study addresses the challenge of distinguishing stable influencers from transient ones among users who are influential at a given point in time. We particularly focus on two distinct types of influencers: source spreaders, who widely disseminate their own content, and brokers, who play a key role in propagating information originating from others. Using six months of retweet data from approximately 19,000 Twitter users, we analyze the characteristics of stable influencers. Our findings reveal that users who have maintained influence in the past are more likely to continue doing so in the future. Furthermore, we develop classification models to predict stable influencers among temporarily influential users, achieving an AUC of approximately 0.89 for source spreaders and 0.81 for brokers. Our experimental results highlight that current influence is a critical factor in classifying influencers, while past influence also significantly contributes, particularly for source spreaders.

cs.SI

Understanding Toxic Interaction Across User and Video Clusters in Social Video Platforms

Social video platforms shape how people access information, while recommendation systems can narrow exposure and increase the risk of toxic interaction. Previous research has often examined text or users in isolation, overlooking the structural context in which such toxic interactions occur. Without considering who interacts with whom and around what content, it is difficult to explain why negative expressions cluster within particular communities. To address this issue, this study focuses on the Chinese social video platform Bilibili, incorporating video-level information as the environment for user expression, modeling users and videos in an interaction matrix. After normalization and dimensionality reduction, we perform separate clustering on both sides of the video-user interaction matrix with K-means. Cluster assignments facilitate comparisons of user behavior, including message length, posting frequency, and source (barrage and comment), as well as textual features such as sentiment and toxicity, and video attributes defined by uploaders. Such a clustering approach integrates structural ties with content signals to identify stable groups of videos and users. We find clear stratification in interaction style (message length, comment ratio) across user clusters, while sentiment and toxicity differences are weak or inconsistent across video clusters. Across video clusters, viewing volume exhibits a clear hierarchy, with higher exposure groups concentrating more toxic expressions. For such a group, platforms should require timely intervention during periods of rapid growth. Across user clusters, comment ratio and message length form distinct hierarchies, and several clusters with longer and comment-oriented messages exhibit lower toxicity. For such groups, platforms should strengthen mechanisms that sustain rational dialogue and encourage engagement across topics.

cs.SI

The Circulate and Recapture Dynamic of Fan Mobility in Agency-Affiliated VTuber Networks

VTuber agencies -- multichannel networks (MCNs) that bundle Virtual YouTubers (VTubers) on YouTube -- curate portfolios of channels and coordinate programming, cross appearances, and branding in the live-streaming VTuber ecosystem. It remains unclear whether affiliation binds fans to a single channel or instead encourages movement within a portfolio that buffers exit, and how these micro level dynamics relate to meso level audience overlap. This study examines how affiliation shapes short horizon viewer trajectories and the organization of audience overlap networks by contrasting agency affiliated and independent VTubers. Using a large, multiyear, fan centered panel of VTuber live stream engagement on YouTube, we construct monthly audience overlap between creators with a similarity measure that is robust to audience size asymmetries. At the micro level, we track retention, changes in the primary creator watched (oshi), and inactivity; at the meso level, we compare structural properties of affiliation specific subgraphs and visualize viewer state transitions. The analysis identifies a pattern of loose mobility: fans tend to remain active while reallocating attention within the same affiliation type, with limited leakage across affiliation type. Network results indicate convergence in global overlap while local neighborhoods within affiliated subgraphs remain persistently denser. Flow diagrams reveal circulate and recapture dynamics that stabilize participation without relying on single channel lock in. We contribute a reusable measurement framework for VTuber live streaming that links micro level trajectories to meso level organization and informs research on creator labor, influencer marketing, and platform governance on video platforms. We do not claim causal effects; the observed regularities are consistent with proximity engineered by VTuber agencies and coordinated recapture.

cs.SI

Global Patterns of Knowledge: Language, Genre, and the Geography of Knowledge

Online platforms, particularly Wikipedia, have become critical infrastructures for providing diverse linguistic and cultural contexts. This human-curated knowledge now forms the foundation for modern AI. However, we have not yet fully explored how knowledge production capability vary across languages and domains. Here, we address this gap by applying economic complexity analysis to understand the editing history of Wikipedia platforms. This approach allows us to infer the latent mode of ``knowledge-production'' of each language community from the diversity and specialization of its contributed content. We reveal that different language communities exhibit distinct specializations, particularly in cultural subjects. Furthermore, we map the global landscape of these production modes, finding that the structure of knowledge production strongly reflects geopolitical boundaries. Our findings suggest that while a common mode of knowledge production exists for standardized topics such as science, it is more diverse for cultural topics or controversial subjects such as conspiracy theories. The association between differences in knowledge production capability and geopolitical factors implies how linguistic and cultural dynamics shape our worldview and the biases embedded in Wikipedia data, a unique, massive, and essential dataset for modern AI.

cs.CY

Comparing User Activity on X and Mastodon

The "Fediverse", a federation of decentralized social media servers, has emerged after a decade in which centralized platforms like X (formerly Twitter) have dominated the landscape. The structure of a federation should affect user activity, as a user selects a server to access the Fediverse and posts are distributed along the structure. This paper reports on the differences in user activity between Twitter and Mastodon, a prominent example of decentralized social media. The target of the analysis is Japanese posts because both Twitter and Mastodon are actively used especially in Japan. Our findings include a larger number of replies on Twitter, more consistent user engagement on mstdn.jp, and different topic preferences on each server.

cs.SI

Prestige bias drives the viral spread of content reposted by influencers in online communities

Cultural evolution theory suggests that prestige bias - whereby individuals preferentially learn from prestigious figures - has played a key role in human ecological success. However, its impact within online environments remains unclear, particularly with respect to whether reposts by prestigious individuals amplify diffusion more effectively than reposts by noninfluential users. We analyzed over 55 million posts and 520 million reposts on Twitter (currently X) to examine whether users with high influence scores (hg indices) more effectively amplified the reach of others' content. Our findings indicate that posts shared by influencers are more likely to be further shared than those shared by non-influencers. This effect persisted over time, especially in viral posts. Moreover, a small group of highly influential users accounted for approximately half of the information flow within repost cascades. These findings demonstrate a prestige bias in information diffusion within the digital society, suggesting that cognitive biases shape content spread through reposting.

cs.SI

Analysis of Psychographic Indicators via LIWC and Their Correlation with CTR for Instagram Ads

The online advertising industry continues to grow and accounts for over 40% of global advertising spending. Online display advertising consists of images and text, and advertisers maximize sales revenue by contacting consumers through advertisements and encouraging them to make purchases. In today's society, where products are becoming more homogenized and needs are diversifying, appealing to consumer psychology through advertisements is becoming increasingly important. However, it is not sufficiently clear what kind of appeal influences consumer psychology. In this study, we quantified the appeal of the text in advertisements for health products and cosmetics, which were actually delivered in Instagram advertisements (one of display advertisements), by applying linguistic inquiry and word count (LIWC). The correlation between click-through rate (CTR) and the text was analyzed. The results showed that negative appeals that arouse consumer anxiety and a sense of crisis were related to CTR.

cs.SI

Comparing Two Counting Methods for Estimating the Probabilities of Strings

There are two methods for counting the number of occurrences of a string in another large string. One is to count the number of places where the string is found. The other is to determine how many pieces of string can be extracted without overlapping. The difference between the two becomes apparent when the string is part of a periodic pattern. This research reports that the difference is significant in estimating the occurrence probability of a pattern. In this study, the strings used in the experiments are approximated from time-series data. The task involves classifying strings by estimating the probability or computing the information quantity. First, the frequencies of all substrings of a string are computed. Each counting method may sometimes produce different frequencies for an identical string. Second, the probability of the most probable segmentation is selected. The probability of the string is the product of all probabilities of substrings in the selected segmentation. The classification results demonstrate that the difference in counting methods is statistically significant, and that the method without overlapping is better.

cs.DS

Follower--Followee Ratio Category and User Vector for Analyzing Following Behavior

Analyzing following behavior is important in many applications. Following behavior may depend on the main intention of the follower. Users may either follow their friends or they may follow celebrities to know more about them. It is difficult to estimate users' intention from their following relationships. In this paper, we propose an approach to analyze following relationships. First, we investigated the similarity between users. Similar followers and followees are likely to be friends. However, when the follower and followee are not similar, it is likely that follower seeks to obtain more information on the followee. Second, we categorized users by the network structure. We then proposed analysis of following behavior based on similarity and category of users estimated from tweets and user data. We confirmed the feasibility of the proposed method through experiments. Finally, we examined users in different categories and analyzed their following behavior.

cs.SI

Analysis of Leading Communities Contributing to arXiv Information Distribution on Twitter

To analyze the impact that arXiv is having on the world, in this paper we propose an arXiv information distribution model on Twitter, which has a three-layer structure: arXiv papers, information spreaders, and information collectors. First, we use the HITS algorithm to analyze the arXiv information diffusion network with users as nodes, which is created from three types of behavior on Twitter regarding arXiv papers: tweeting, retweeting, and liking. Next, we extract communities from the network of information spreaders with positive authority and hub degrees using the Louvain method, and analyze the relationship and roles of information spreaders in communities using research field, linguistic, and temporal characteristics. From our analysis using the tweet and arXiv datasets, we found that information about arXiv papers circulates on Twitter from information spreaders to information collectors, and that multiple communities of information spreaders are formed according to their research fields. It was also found that different communities were formed in the same research field, depending on the research or cultural background of the information spreaders. We were able to identify two types of key persons: information spreaders who lead the relevant field in the international community and information spreaders who bridge the regional and international communities using English and their native language. In addition, we found that it takes some time to gain trust as an information spreader.

cs.DL

Do you trust experts on Twitter?: Successful correction of COVID-19-related misinformation

This study focuses on how scientifically-correct information is disseminated through social media, and how misinformation can be corrected. We have identified examples on Twitter where scientific terms that have been misused have been rectified and replaced by scientifically-correct terms through the interaction of users. The results show that the percentage of correct terms ("variant" or "COVID-19 variant") being used instead of the incorrect terms ("strain") on Twitter has already increased since the end of December 2020. This was about a month before the release of an official statement by the Japanese Association for Infectious Diseases regarding the correct terminology, and the use of terms on social media was faster than it was in television. Some Twitter users who quickly started using the correct term were more likely to retweet messages sent by leading influencers on Twitter, rather than messages sent by traditional media or portal sites. However, a few Twitter users continued to use wrong terms even after March 2021, even though the use of the correct terms was widespread. Further analysis of their tweets revealed that they were quoting sources that differed from that of other users. This study empirically verified that self-correction occurs even on Twitter, which is often known as a "hotbed for spreading rumors." The results of this study also suggest that influencers with expertise can influence the direction of public opinion on social media and that the media that users usually cite can also affect the possibility of behavioral changes.

cs.SI

Feature Selective Likelihood Ratio Estimator for Low- and Zero-frequency N-grams

In natural language processing (NLP), the likelihood ratios (LRs) of N-grams are often estimated from the frequency information. However, a corpus contains only a fraction of the possible N-grams, and most of them occur infrequently. Hence, we desire an LR estimator for low- and zero-frequency N-grams. One way to achieve this is to decompose the N-grams into discrete values, such as letters and words, and take the product of the LRs for the values. However, because this method deals with a large number of discrete values, the running time and memory usage for estimation are problematic. Moreover, use of unnecessary discrete values causes deterioration of the estimation accuracy. Therefore, this paper proposes combining the aforementioned method with the feature selection method used in document classification, and shows that our estimator provides effective and efficient estimation results for low- and zero-frequency N-grams.

cs.CL

Comparison of Indicators of Location Homophily Using Twitter Follow Graph

Location homophily is a tendency of Twitter users whose followers tend to be in the same or nearby areas. Intuitively, although users with a higher number of follower relationships might have negative homophily indicators, it is worth consulting actual Twitter data. Moreover, there may be certain functions regarding the numbers of friends and followers that are more directly correlated to the homophily. In this study, the ratio of the number of friends to the number of followers is shown to be a more effective negative indicator of homophily, and the results for 10 different countries are verified.

cs.SI

Unified Likelihood Ratio Estimation for High- to Zero-frequency N-grams

Likelihood ratios (LRs), which are commonly used for probabilistic data processing, are often estimated based on the frequency counts of individual elements obtained from samples. In natural language processing, an element can be a continuous sequence of $N$ items, called an $N$-gram, in which each item is a word, letter, etc. In this paper, we attempt to estimate LRs based on $N$-gram frequency information. A naive estimation approach that uses only $N$-gram frequencies is sensitive to low-frequency (rare) $N$-grams and not applicable to zero-frequency (unobserved) $N$-grams; these are known as the low- and zero-frequency problems, respectively. To address these problems, we propose a method for decomposing $N$-grams into item units and then applying their frequencies along with the original $N$-gram frequencies. Our method can obtain the estimates of unobserved $N$-grams by using the unit frequencies. Although using only unit frequencies ignores dependencies between items, our method takes advantage of the fact that certain items often co-occur in practice and therefore maintains their dependencies by using the relevant $N$-gram frequencies. We also introduce a regularization to achieve robust estimation for rare $N$-grams. Our experimental results demonstrate that our method is effective at solving both problems and can effectively control dependencies.

cs.CL

Corrective Information Does Not Necessarily Curb Social Disruption

The spread of misinformation can cause social confusion. The authenticity of information on a social networking service (SNS) is unknown, and false information can be easily spread. Consequently, many studies have been conducted on methods to control the spread of misinformation on social networking sites. However, few studies have examined the impact of the spread of misinformation and its corrections on society. This study models the impact of the reduction of misinformation and the diffusion of corrective information on social disruption, and it identifies the features of this impact. In this study, we analyzed misinformation regarding the shortage of toilet paper during the 2020 COVID-19 epidemic, its corrections, and the excessive purchasing caused by this information. First, we analyze the amount of misinformation and corrective information spread on SNS, and we create a regression model to estimate the real-world impact of misinformation and its correction. This model is used to analyze the change in real-world impact corresponding to the change in the diffusion of misinformation and corrective information. Our analysis shows that the corrective information was spread to a much greater extent than the misinformation. In addition, our model reveals that the corrective information was what caused the excessive purchasing behavior. As a result of our further analysis, we found that the amount of diffusion of corrective information required to minimize the impact on the real world depends on the amount of the diffusion of misinformation.

cs.SI