SearcharxivSearch

arXiv subjects

Hana Matatov

Publications and source records attributed to Hana Matatov.

4 recordsLinked to original sources

Examining the Prevalence and Dynamics of AI-Generated Media in Art Subreddits

Broadly accessible generative AI models like Dall-E have made it possible for anyone to create compelling visual art. In online communities, the introduction of AI-generated content (AIGC) may impact social dynamics, for example causing changes in who is posting content, or shifting the norms or the discussions around the posted content if posts are suspected of being generated by AI. We take steps towards examining the potential impact of AIGC on art-related communities on Reddit. We distinguish between communities that disallow AI content and those without such a direct policy. We look at image-based posts in these communities where the author transparently shares that the image was created by AI, and at comments in these communities that suspect or accuse authors of using generative AI. We find that AI posts (and accusations) have played a surprisingly small part in these communities through the end of 2023, accounting for fewer than 0.5% of the image-based posts. However, even as the absolute number of author-labeled AI posts dwindles over time, accusations of AI use remain more persistent. We show that AI content is more readily used by newcomers and may help increase participation if it aligns with community rules. However, the tone of comments suspecting AI use by others has become more negative over time, especially in communities that do not have explicit rules about AI. Overall, the results show the changing norms and interactions around AIGC in online communities designated for creativity.

cs.AI

Stop the [Image] Steal: The Role and Dynamics of Visual Content in the 2020 U.S. Election Misinformation Campaign

Images are powerful. Visual information can attract attention, improve persuasion, trigger stronger emotions, and is easy to share and spread. We examine the characteristics of the popular images shared on Twitter as part of "Stop the Steal", the widespread misinformation campaign during the 2020 U.S. election. We analyze the spread of the forty most popular images shared on Twitter as part of this campaign. Using a coding process, we categorize and label the images according to their type, content, origin, and role, and perform a mixed-method analysis of these images' spread on Twitter. Our results show that popular images include both photographs and text rendered as image. Only very few of these popular images included alleged photographic evidence of fraud; and none of the popular photographs had been manipulated. Most images reached a significant portion of their total spread within several hours from their first appearance, and both popular- and less-popular accounts were involved in various stages of their spread.

cs.SI

Dataset and Case Studies for Visual Near-Duplicates Detection in the Context of Social Media

The massive spread of visual content through the web and social media poses both challenges and opportunities. Tracking visually-similar content is an important task for studying and analyzing social phenomena related to the spread of such content. In this paper, we address this need by building a dataset of social media images and evaluating visual near-duplicates retrieval methods based on image retrieval and several advanced visual feature extraction methods. We evaluate the methods using a large-scale dataset of images we crawl from social media and their manipulated versions we generated, presenting promising results in terms of recall. We demonstrate the potential of this method in two case studies: one that shows the value of creating systems supporting manual content review, and another that demonstrates the usefulness of automatic large-scale data analysis.

cs.IR

VoterFraud2020: a Multi-modal Dataset of Election Fraud Claims on Twitter

The wide spread of unfounded election fraud claims surrounding the U.S. 2020 election had resulted in undermining of trust in the election, culminating in violence inside the U.S. capitol. Under these circumstances, it is critical to understand the discussions surrounding these claims on Twitter, a major platform where the claims were disseminated. To this end, we collected and released the VoterFraud2020 dataset, a multi-modal dataset with 7.6M tweets and 25.6M retweets from 2.6M users related to voter fraud claims. To make this data immediately useful for a diverse set of research projects, we further enhance the data with cluster labels computed from the retweet graph, each user's suspension status, and the perceptual hashes of tweeted images. The dataset also includes aggregate data for all external links and YouTube videos that appear in the tweets. Preliminary analyses of the data show that Twitter's user suspension actions mostly affected a specific community of voter fraud claim promoters, and exposes the most common URLs, images and YouTube videos shared in the data.

cs.SI