SearcharxivSearch

arXiv subjects

Jasmin Wyss

Publications and source records attributed to Jasmin Wyss.

4 recordsLinked to original sources

Stop Abandoning Me: Exploring the Landscape of Unmaintained Intimate Partner Abuse Support Applications

Many support applications are developed to help users living through intimate partner abuse (IPA). However, many of those projects get abandoned, because the application was a prototype that never got turned into an actual product, funding ran out, or the people maintaining it moved onto other projects. This abandonment can have devastating consequences because the users of these applications are vulnerable by definition. In this ongoing research, we aim to measure the abandonment of intimate partner abuse support applications. Our preliminary results with a dataset of 197 support tools indicate that 58.9\% of applications have lost support over time, either no longer receiving updates or not being available online at all.

cs.CY

It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance

Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.

cs.CY

Traces of Abuse: How Generative AI Impacts Image-Based Sexual Abuse (IBSA) Investigations

The introduction of generative AI (GAI) into the workflow of image-based sexual abuse (IBSA) only worsened the ease of creation and distribution, victimizing more people than ever. We outline how the introduction of generative AI (GAI-IBSA) impacts the creation of traces and the type of reasoning they allow. We illustrate the impact by comparing the forensic traces available in four different IBSA scenarios. We discuss the impacts on the (possibility of) investigation, arguing that the advent of generative AI overall benefits abusers by making perpetration easier, and the perpetrator harder to trace.

cs.CY

Unfair Mistakes on Social Media: How Demographic Characteristics influence Authorship Attribution

Authorship attribution techniques are increasingly being used in online contexts such as sock puppet detection, malicious account linking, and cross-platform account linking. Yet, it is unknown whether these models perform equitably across different demographic groups. Bias in such techniques could lead to false accusations, account banning, and privacy violations disproportionately impacting users from certain demographics. In this paper, we systematically audit authorship attribution for bias with respect to gender, native language, and age. We evaluate fairness in 3 ways. First, we evaluate how the proportion of users with a certain demographic characteristic impacts the overall classifier performance. Second, we evaluate if a user's demographic characteristics influence the probability that their texts are misclassified. Our analysis indicates that authorship attribution does not demonstrate bias across demographic groups in the closed-world setting. Third, we evaluate the types of errors that occur when the true author is removed from the suspect set, thereby forcing the classifier to choose an incorrect author. Unlike the first two settings, this analysis demonstrates a tendency to attribute authorship to users who share the same demographic characteristic as the true author. Crucially, these errors do not only include texts that deviate from a user's usual style, but also those that are very close to the author's average. Our results highlight that though a model may appear fair in the closed-world setting for a performant classifier, this does not guarantee fairness when errors are inevitable.

cs.SI