arXiv · 2601.14124
Style Transfer as Bias Mitigation: Diffusion Models for Synthetic Mental Health Text for Arabic
Abstract
Synthetic data offers a promising solution for mitigating data scarcity and demographic bias in mental health analysis, yet existing approaches largely rely on pretrained large language models (LLMs), which may suffer from limited output diversity and propagate biases inherited from their training data. In this work, we propose a pretraining-free diffusion-based approach for synthetic text generation that frames bias mitigation as a style transfer problem. Using the CARMA Arabic mental health corpus, which exhibits a substantial gender imbalance, we focus on male-to-female style transfer to augment underrepresented female-authored content. We construct five datasets capturing varying linguistic and semantic aspects of gender expression in Arabic and train separate diffusion models for each setting. Quantitative evaluations demonstrate consistently high semantic fidelity between source and generated text, alongside meaningful surface-level stylistic divergence, while qualitative analysis confirms linguistically plausible gender transformations. Our results show that diffusion-based style transfer can generate high-entropy, semantically faithful synthetic data without reliance on pretrained LLMs, providing an effective and flexible framework for mitigating gender bias in sensitive, low-resource mental health domains.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Saad Mankarious, Aya Zirikly. 2026-01-20. Style Transfer as Bias Mitigation: Diffusion Models for Synthetic Mental Health Text for Arabic. https://arxiv.org/abs/2601.14124
Cite the original work for its findings. Save a collection to share your selection of sources.