SearcharxivSearch

arXiv subjects

Yasir Zaki

Publications and source records attributed to Yasir Zaki.

At least 19 recordsLinked to original sources

Understanding AI Provider Recommendations in Local Service Markets

When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.

cs.CY

Informational Help-Seeking on Reddit Did Not Decline After ChatGPT

Did people stop asking other people for advice online once generative AI could answer their questions? Prior work on ChatGPT's effect on online help-seeking disagrees in both size and sign, in part because no study has compared affected communities against similar communities that AI cannot easily substitute for, over the same months. In this paper, we track monthly post counts in 26 Reddit informational communities against 90 size-comparable hobby communities over the same six calendar months before and after the launch of ChatGPT. We also repeat the entire analysis at 66 earlier dates, before ChatGPT existed, to see what our method reports when no ChatGPT-effect exists. We find that informational help-seeking did not decline. Our results rule out any decline in posting larger than 3.4%, far smaller than the 8% to 25% declines documented in prior work. Steady post counts could still be misleading if AI-written posts had replaced human ones. We test this possibility by scoring 274,411 posts and 223,775 comments with AI-text detectors, compared in a way that cancels out detector false-positives on human-written text. AI-written posts rose only 2-3 percentage points more in informational communities than in hobby communities, short of the 5.1 points that would be needed to hide even the smallest decline previously reported for Reddit. In addition, the comments people receive show no such rise at all. Why, then, do published studies disagree? Reddit community types were already drifting apart before ChatGPT existed, at rates comparable to every published estimate, and without same-time controls, that drift can look like an effect of generative AI. Our own largest estimate, an 18% fall in posts to low-stakes curiosity communities, matches its pre-existing trend. Humans still ask humans for help, and, as far as detection can tell, humans still answer them.

cs.SI

Chance, Persistent Advantage, and the Generative-AI Era in Open-Source Package Careers

Studies of careers in science, film, music, and books report a common pattern. When a person's most successful work arrives is close to a random draw over the works they produce. How large their successes tend to be, in contrast, follows a stable, person-specific factor. We test whether this pattern holds for open-source software careers and whether it changed when generative AI coding tools arrived. From the complete public record of GitHub push events (2015-2025), we reconstruct 102.2M career works by 6.15M contributors, and for the 908k contributors whose repositories publish packages, we measure each work's impact by how many downstream packages come to depend on it. First, we find that the timing of a career's biggest hit is close to a lottery over their works, as in science and the arts, with a small, replicable lean toward early career that grows as careers get longer. Second, some coders reliably produce higher-impact work than others, but this lasting personal factor accounts for only part of why impact persists (about a fifth in our primary specification); the rest behaves like momentum, success feeding on itself for a period of time. Third, within the same contributors, this structure did not change after ChatGPT's release. The stable factor's weight grew by about as much as it grew for an earlier cohort that simply aged, and subtracting the effect of aging from the effect of generative AI puts the shift at +0.03 (95% CI [-0.22, +0.23]), indistinguishable from zero. The success pattern documented in science and the arts therefore describes open-source careers too, and it shows no detectable break across the arrival of generative AI. These results have implications for how track records on open platforms should be read and on what to expect from generative AI for the careers built on them.

cs.SI

Price Dislocations, News Citations, and Epistemic Leverage on Polymarket

Prediction-market probabilities increasingly appear in news coverage, yet little is known about which market movements become news or how much trading money sits behind the numbers journalists quote. Unlike a poll, a market price can be moved by anyone willing to trade, so the cost of manufacturing a number that circulates as news bears directly on the information environment. We link 173.7 million signed Polymarket trades to news coverage from 2024-2025. From 6,990 articles mentioning prediction-market venues, an LLM-based, human-validated matcher extracts 1,582 sentences quoting market odds and attributes 918 to the specific market whose price they cite. We then detect 44,976 price dislocations, movements of at least five percentage points backed by concentrated one-sided trading, and ask whether a market is cited more often afterward. In the days after a dislocation, a market's citation rate is about 33% higher than its matched baseline (log citation-rate ratio $τ_{\mathrm{cite}}=0.283$, permutation $p=0.001$), robust to binary and Poisson count outcomes. Yet move size is not the strongest predictor of citation: prominence dominates (standardized $β=0.610$ vs. $β=0.159$ for move size). Finally, we combine the dollar flow behind a given price change with observed citation rates into a metric we call epistemic leverage, the dollars needed to move a market five points and have the move cited. It stays near \$0.7-1.0 million across prominence quintiles, because cheaper-to-move markets are proportionally less likely to be cited. The implied threat model centers not on the long tail of cheaply moved markets but on the few prominent markets newsrooms treat as informational infrastructure, where a seven-figure price of influence sits within the budgets of actors with a large stake in the quoted number. We release aggregate event-study data and validation materials.

physics.soc-ph

The Political Ideology of Large Language Models: Measurement, Inconsistency, and Persuasive Influence

Large Language Models (LLMs) are a transformational technology, fundamentally changing how people obtain information and interact with the world. As people become increasingly reliant on them for an enormous variety of tasks, a body of academic research has developed to examine these models for inherent biases, especially political biases, often finding them small. We challenge this prevailing wisdom. First, by comparing 43 LLMs to legislators, judges, and a nationally representative sample of U.S. voters, we show that LLMs' apparently moderate overall partisan positioning is the net result of offsetting strongly partisan expressed positions on specific topics, much like moderate voters. Second, in a pre-registered randomized experiment, we show that LLMs can exert persuasive influence on political attitudes. Voters randomized to discuss a policy issue with an LLM shift toward that model's measured ideological position by 3.5 percentage points on average, an effect at least as large as those produced by professional campaign advertising. Explicitly prompting a model to argue one side of the issue shifts attitudes by more than 10 percentage points relative to unsteered conversations, and this steering accounts for the pooled effect. When the same models converse naturally, without steering, we detect no persuasive effect, and our confidence interval rules out effects as small as the pre-registered smallest effect of interest. Contrary to expectations, these persuasive effects are not moderated by familiarity with LLMs, news consumption, or interest in politics. LLMs, especially those controlled by private companies or governments, may become a powerful and targeted vector for political influence.

cs.CY

Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?

We study peer review at ICLR, a large machine-learning conference whose complete review record, including rejected submissions, is public. Reviewers score each submission; for the borderline band whose scores do not settle an outcome, an area chair makes a discretionary accept-or-reject call. We ask whether that call is even-handed: do authors from prestigious institutions, WEIRD countries, or all-male teams get the benefit of the doubt at the margin? Across ICLR 2019-2025 (31,711 submissions; 10,416 borderline), borderline papers without a top-25-institution author are accepted at a 0.5 to 1.6 percentage point lower rate at the same reviewer scores. The gap arises at the discretionary stage, reappears out-of-sample in the pre-registered ICLR 2026 cohort, and concentrates almost entirely among submissions identifiable through a pre-decision arXiv preprint (-3.4 vs. -0.2 points). Equal scores need not mean equal papers: an area chair may respond to quality the scores miss. We apply a robust outcome test, which concludes discrimination only when the group accepted at a lower rate also realizes better downstream outcomes. We measure five outcomes (citations, disruption, two forms of novelty, eventual venue) on both sides of the decision, including the first "ones that got away" test of rejected submissions. Our headline result is a null: across a pre-registered family of 27 tests, no disparity concordant with the decision-rate gap survives correction; we find no evidence that any group faced a higher bar on the outcomes we measure. That null is not an exoneration. A pre-decision preprint pierces the blind through policy-permitted means, and the acceptance gap lives almost entirely in that porosity, consistent with area chairs using revealed institutional prestige as a prior: statistical discrimination that outcome tests may not detect, and a practice double-blind review exists to prevent.

cs.DL

On-Screen Inertia: Persistent Racial and Gender Disparities in Hollywood Film (1900-2024)

Hollywood has diversified its casts. Whether this has translated into structural change in how those actors are positioned within narratives remains largely unexamined. Drawing on 76,815 U.S. English-language films (1900-2024) and over 3.1 million cast and crew entries, we move beyond headcounts to examine long-term inclusion trends through network centrality, occupational stereotypes, crew-to-cast diversity pathways, and financial outcomes. We find evidence of what we term on-screen inertia. While the raw inclusion of women and racial minorities has increased modestly, White actors have become more overrepresented relative to the U.S. Census in recent decades, not less. Within the visibility layer, women face a consistent longevity penalty with significantly shorter careers than men, and visual depictions framing men as dominant and women as sensual have remained stable since the 1950s. Structurally, White actors retain disproportionate network centrality; women achieve parity in centrality and lead billing yet cluster in secondary co-lead roles; and occupational stereotypes anchoring racial and gender groups to specific labor categories persist largely unchanged across the pre- and post-2000 periods. Crew diversity associates with cast inclusion only along matching demographic lines (i.e., racial with racial, gender with gender) and does not extend to narrative centrality, revealing a structural ceiling on hiring-based interventions. Critically, we find no consistent market penalty for diversity across decades of box office returns and audience ratings, eliminating the primary rationalization for these practices. Together, these findings demonstrate that Hollywood's representational inequalities are not a rational market response, but are an institutionally sustained choice.

cs.CY

Causal evidence of racial and institutional biases in accessing paywalled articles and scientific data

Scientific progress depends on researchers' ability to access and build upon the work of others. Yet, much published work remains behind expensive paywalls, and even accessible articles often rest on datasets shared only "upon reasonable request" to the authors. Researchers can try to overcome these barriers through informal channels, such as emailing authors directly, but whether such channels are hindered by racial or institutional biases remains unknown. Here we combine survey data, semi-structured interviews, large-scale observational analysis, and two randomized audit experiments to examine disparities in access to scientific knowledge. Surveyed researchers in the Global South report markedly lower institutional access to the literature and depend more heavily on informal channels to obtain papers and data; interviews elaborate the workarounds and racialized frictions they encounter. Our analysis of 250 million articles reveals that Global South researchers cite paywalled papers at significantly lower rates than Global North counterparts--a gap associated with reduced knowledge breadth and scholarly impact. Using citation-context classification, we further find that papers whose data is available only upon request are less likely to be cited for reusing their data, a penalty falling disproportionately on the Global South. To probe mechanisms, we conduct two email audit studies in which fictional PhD students differing in racial background and institutional affiliation request paywalled articles (N = 18,000) and datasets (N = 16,000). Racial identity influences response rates to both requests, whereas institutional affiliation influences access to datasets. These findings reveal how informal gatekeeping can perpetuate structural inequities in science, highlighting the need for stronger data-sharing mandates and more equitable open-access policies.

cs.DL

Schadenfreude in the Digital Public Sphere: A cross-national and decade-long analysis of Facebook news engagement

Schadenfreude, or the pleasure derived from others' misfortunes, has become a visible and performative feature of online news engagement, yet little is known about its prevalence, dynamics, or social patterning. We examine schadenfreude on Facebook over a ten-year period across nine major news publishers in the United States, the United Kingdom, and India (one left-leaning, one right-leaning, and one centrist per country). Using a combination of human annotation and machine-learning classification, we identify posts describing misfortune and detect schadenfreude in nearly one million associated comments. We find that while sadness and anger dominate reactions to misfortune posts, laughter and amusement form a substantial and patterned minority. Schadenfreude is most frequent in moralized and political contexts, higher among right-leaning audiences, and more pronounced in India than in the United States or United Kingdom. Temporal and regression analyses further reveal asymmetric relationships between political power and schadenfreude: left-leaning outlets display "power-licensed" schadenfreude that increases when their party governs, while right-leaning outlets exhibit "power-compensatory" schadenfreude that intensifies in opposition. Together, our findings move beyond anecdotal accounts to map schadenfreude as a dynamic, context-dependent feature of digital discourse, revealing how it evolves over time and across ideological and cultural divides.

cs.SI

(Mis-)Informed Consent: Predatory Apps and the Exploitation of Populations with Limited Literacy

Among populations with limited literacy in emerging digital markets, the adoption of mobile phones, combined with comprehension barriers and poor cybersecurity hygiene, has created hidden privacy risks. This paper examines how informed consent is often abused by predatory financial applications, leading to financial scams that disproportionately affect users with low literacy. We focus on predatory loan, gambling, and trading apps, analyzing a dataset of 50 Google Play Store apps to measure how many omit or obfuscate critical privacy disclosures. We also evaluate comprehension gaps among users with low literacy via a targeted user study and assess whether Large Language Model (LLM)-generated summaries, translations, and visual cues can improve consent clarity. Our findings show that 85% of study participants did not understand basic app permissions, underscoring the urgent need for stronger regulatory oversight and scalable LLM-driven privacy-literacy tools.

cs.CY

Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks

The rapid spread of misinformation on digital platforms threatens public discourse, emotional stability, and decision-making. While prior work has explored various adversarial attacks in misinformation detection, the specific transformations examined in this paper have not been systematically studied. In particular, we investigate language-switching across English, French, Spanish, Arabic, Hindi, and Chinese, followed by translation. We also study query length inflation preceding summarization and structural reformatting into multiple-choice questions. In this paper, we present a multilingual, multi-agent large language model framework with retrieval-augmented generation that can be deployed as a web plugin into online platforms. Our work underscores the importance of AI-driven misinformation detection in safeguarding online factual integrity against diverse attacks, while showcasing the feasibility of plugin-based deployment for real-world web applications.

cs.CL

Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes

The rapid adoption of generative AI tools is reshaping how scholars produce and communicate knowledge, raising questions about who benefits and who is left behind. We analyze over 230,000 Scopus-indexed computer science articles between 2021 and 2025 to examine how AI-assisted writing alters scholarly visibility across regions. Using zero-shot detection of AI-likeness, we track stylistic changes in writing and link them to citation counts, journal placement, and global citation flows before and after ChatGPT. Our findings reveal uneven outcomes: authors in the Global East adopt AI tools more aggressively, yet Western authors gain more per unit of adoption due to pre-existing penalties for "humanlike" writing. Prestigious journals continue to privilege more human-sounding texts, creating tensions between visibility and gatekeeping. Network analyses show modest increases in Eastern visibility and tighter intra-regional clustering, but little structural integration overall. These results highlight how AI adoption reconfigures the labor of academic writing and reshapes opportunities for recognition.

cs.CY

Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Perspective

English is the predominant language on the web, powering nearly half of the world's top ten million websites. Support for multilingual content is nevertheless growing, with many websites increasingly combining English with regional or native languages in both visible content and hidden metadata. This multilingualism introduces significant barriers for users with visual impairments, as assistive technologies like screen readers frequently lack robust support for non-Latin scripts and misrender or mispronounce non-English text, compounding accessibility challenges across diverse linguistic contexts. Yet, large-scale studies of this issue have been limited by the lack of comprehensive datasets on multilingual web content. To address this gap, we introduce LangCrUX, the first large-scale dataset of 120,000 popular websites across 12 languages that primarily use non-Latin scripts. Leveraging this dataset, we conduct a systematic analysis of multilingual web accessibility and uncover widespread neglect of accessibility hints. We find that these hints often fail to reflect the language diversity of visible content, reducing the effectiveness of screen readers and limiting web accessibility. We finally propose Kizuki, a language-aware automated accessibility testing extension to account for the limited utility of language-inconsistent accessibility hints.

cs.CL

Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models

The rise of social media and online communication platforms has led to the spread of Arabic textual posts and memes as a key form of digital expression. While these contents can be humorous and informative, they are also increasingly being used to spread offensive language and hate speech. Consequently, there is a growing demand for precise analysis of content in Arabic text and memes. This paper explores the potential of large language models to effectively identify hope, hate speech, offensive language, and emotional expressions within such content. We evaluate the performance of base LLMs, fine-tuned LLMs, and pre-trained embedding models. The evaluation is conducted using a dataset of Arabic textual speech and memes proposed in the ArabicNLP MAHED 2025 challenge. The results underscore the capacity of LLMs such as GPT-4o-mini, fine-tuned with Arabic textual speech, and Gemini Flash 2.5, fine-tuned with Arabic memes, to deliver the superior performance. They achieve up to 72.1%, 57.8%, and 79.6% macro F1 scores for tasks 1, 2, and 3, respectively, and secure first place overall in the Mahed 2025 challenge. The proposed solutions offer a more nuanced understanding of both text and memes for accurate and efficient Arabic content moderation systems.

cs.CL

Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases

Islamic inheritance domain holds significant importance for Muslims to ensure fair distribution of shares between heirs. Manual calculation of shares under numerous scenarios is complex, time-consuming, and error-prone. Recent advancements in Large Language Models (LLMs) have sparked interest in their potential to assist with complex legal reasoning tasks. This study evaluates the reasoning capabilities of state-of-the-art LLMs to interpret and apply Islamic inheritance laws. We utilized the dataset proposed in the ArabicNLP QIAS 2025 challenge, which includes inheritance case scenarios given in Arabic and derived from Islamic legal sources. Various base and fine-tuned models, are assessed on their ability to accurately identify heirs, compute shares, and justify their reasoning in alignment with Islamic legal principles. Our analysis reveals that the proposed majority voting solution, leveraging three base models (Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3), outperforms all other models that we utilized across every difficulty level. It achieves up to 92.7% accuracy and secures the third place overall in Task 1 of the Qias 2025 challenge.

cs.CL

Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks

Recent progress in large language models (LLMs) has showcased impressive proficiency in numerous Arabic natural language processing (NLP) applications. Nevertheless, their effectiveness in Arabic medical NLP domains has received limited investigation. This research examines the degree to which state-of-the-art LLMs demonstrate and articulate healthcare knowledge in Arabic, assessing their capabilities across a varied array of Arabic medical tasks. We benchmark several LLMs using a medical dataset proposed in the Arabic NLP AraHealthQA challenge in MedArabiQ2025 track. Various base LLMs were assessed on their ability to accurately provide correct answers from existing choices in multiple-choice questions (MCQs) and fill-in-the-blank scenarios. Additionally, we evaluated the capacity of LLMs in answering open-ended questions aligned with expert answers. Our results reveal significant variations in correct answer prediction accuracy and low variations in semantic alignment of generated answers, highlighting both the potential and limitations of current LLMs in Arabic clinical contexts. Our analysis shows that for MCQs task, the proposed majority voting solution, leveraging three base models (Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3), outperforms others, achieving up to 77% accuracy and securing first place overall in the Arahealthqa 2025 shared task-track 2 (sub-task 1) challenge. Moreover, for the open-ended questions task, several LLMs were able to demonstrate excellent performance in terms of semantic alignment and achieve a maximum BERTScore of 86.44%.

cs.CL

Towards Next Generation Immersive Applications in 5G Environments

The Multi-user Immersive Reality (MIR) landscape is evolving rapidly, with applications spanning virtual collaboration, entertainment, and training. However, wireless network limitations create a critical bottleneck, struggling to meet the high-bandwidth and ultra-low latency demands essential for next-generation MIR experiences. This paper presents Hera, a modular framework for next-generation immersive applications, comprising a high-level streaming and synchronization layer for AR/VR systems and a low-level delay-based QoE-aware rate control protocol optimized for dynamic wireless environments. The Hera framework integrates application-aware streaming logic with a QoE-centric rate control core, enabling adaptive video quality, multi-user fairness, and low-latency communication across challenging 5G network conditions. We demonstrate that Hera outperforms existing state-of-the-art rate control algorithms by maintaining up to 66% lower latencies with comparable throughput performance, higher visual quality with 50% average bitrate improvements in our analysis, and improved fairness. By bridging the gap between application-level responsiveness and network-level adaptability, Hera lays the foundation for more scalable, robust, and high-fidelity multi-user immersive experiences.

cs.NI

Self-Regulating Cars: Automating Traffic Control in Free Flow Road Networks

Free-flow road networks, such as suburban highways, are increasingly experiencing traffic congestion due to growing commuter inflow and limited infrastructure. Traditional control mechanisms, such as traffic signals or local heuristics, are ineffective or infeasible in these high-speed, signal-free environments. We introduce self-regulating cars, a reinforcement learning-based traffic control protocol that dynamically modulates vehicle speeds to optimize throughput and prevent congestion, without requiring new physical infrastructure. Our approach integrates classical traffic flow theory, gap acceptance models, and microscopic simulation into a physics-informed RL framework. By abstracting roads into super-segments, the agent captures emergent flow dynamics and learns robust speed modulation policies from instantaneous traffic observations. Evaluated in the high-fidelity PTV Vissim simulator on a real-world highway network, our method improves total throughput by 5%, reduces average delay by 13%, and decreases total stops by 3% compared to the no-control setting. It also achieves smoother, congestion-resistant flow while generalizing across varied traffic patterns, demonstrating its potential for scalable, ML-driven traffic management.

cs.LG