SearcharxivSearch

arXiv subjects

Heger Arfaoui

Publications and source records attributed to Heger Arfaoui.

3 recordsLinked to original sources

Phase Boundary of a Stochastic Watts-Threshold SIS Model on Random Networks

Complex contagion models, in which adoption requires reinforcement from multiple neighbors, have been extensively studied in the monotone (no-recovery) setting, but the phase diagram of threshold models with SIS-like recovery on networks remains unmapped. We study a stochastic Watts-threshold SIS model on Erdos-Renyi and Barabasi-Albert networks and reconstruct its extinction-persistence phase boundary in the joint parameter space of transmission rate $\beta$, adoption threshold $\theta$, and infectious duration $d$. Using adaptive Delaunay-based sampling and weighted logistic regression on over 180,000 Monte Carlo trials, we find that: (i) the boundary is well described by a six-parameter interaction model whose structure is invariant across both topologies; (ii) the transition is sharp, with the 10-90\% extinction-probability band spanning only $\Delta\theta \approx 0.005$-$0.008$; and (iii) the adoption threshold is the dominant parameter governing epidemic feasibility, with transmission rate and infectious duration playing secondary and asymmetric roles. The characterization provides a quantitative reference for the complex-contagion analogue of the classical SIS epidemic threshold.

cs.SI

A Reproducible Framework for Neural Topic Modeling in Focus Group Analysis

Focus group discussions generate rich qualitative data but their analysis traditionally relies on labor-intensive manual coding that limits scalability and reproducibility. We present a systematic framework for applying BERTopic to focus group transcripts using data from ten focus groups exploring HPV vaccine perceptions in Tunisia (1,075 utterances). We conducted comprehensive hyperparameter exploration across 27 configurations, evaluating each through bootstrap stability analysis, performance metrics, and comparison with LDA baseline. Bootstrap analysis revealed that stability metrics (NMI and ARI) exhibited strong disagreement (r = -0.691) and showed divergent relationships with coherence, demonstrating that stability is multifaceted rather than monolithic. Our multi-criteria selection framework yielded a 7-topic model achieving 18\% higher coherence than optimized LDA (0.573 vs. 0.486) with interpretable topics validated through independent human evaluation (ICC = 0.700, weighted Cohen's kappa = 0.678). These findings demonstrate that transformer-based topic modeling can extract interpretable themes from small focus group transcript corpora when systematically configured and validated, while revealing that quality metrics capture distinct, sometimes conflicting constructs requiring multi-criteria evaluation. We provide complete documentation and code to support reproducibility.

cs.CL

TEET! Tunisian Dataset for Toxic Speech Detection

The complete freedom of expression in social media has its costs especially in spreading harmful and abusive content that may induce people to act accordingly. Therefore, the need of detecting automatically such a content becomes an urgent task that will help and enhance the efficiency in limiting this toxic spread. Compared to other Arabic dialects which are mostly based on MSA, the Tunisian dialect is a combination of many other languages like MSA, Tamazight, Italian and French. Because of its rich language, dealing with NLP problems can be challenging due to the lack of large annotated datasets. In this paper we are introducing a new annotated dataset composed of approximately 10k of comments. We provide an in-depth exploration of its vocabulary through feature engineering approaches as well as the results of the classification performance of machine learning classifiers like NB and SVM and deep learning models such as ARBERT, MARBERT and XLM-R.

cs.CL