SearcharxivSearch

arXiv subjects

Muhammad Atif Qureshi

Publications and source records attributed to Muhammad Atif Qureshi.

2 recordsLinked to original sources

Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation.

cs.CL

An Analysis of Privacy-Aware Personalization Signals by Using Online Evaluation Methods

Personalization despite being an effective solution to the problem information overload remains tricky on account of multiple dimensions to consider. Furthermore, the challenge of avoiding overdoing personalization involves estimation of a user's preferences in relation to different queries. This work is an attempt to make inferences about when personalization would be beneficial by relating observable user behavior to his/her social network usage patterns and user-generated content. User behavior on a search system is observed by means of team-draft interleaving whereby results from two retrieval functions are presented in an interleaved manner, and user clicks are utilised to infer preference for a certain retrieval function. This improves upon earlier work which had limited usefulness due to reliance on user survey results; our findings may aid real-time personalization in search systems by detecting a user-related and query-related personalization signals.

cs.IR