SearcharxivSearch

arXiv subjects

Melanie Jouaiti

Publications and source records attributed to Melanie Jouaiti.

3 recordsLinked to original sources

All I Hear is Noise: Investigating Clever Hans Effects in Clinical Speech Datasets

Recent work revealed a striking Clever Hans effect in the Pitt dataset, where Alzheimer's detection achieved nearly 100% accuracy using only silent audio segments. This raises serious concerns about hidden confounding factors in speech-based health datasets. We systematically investigate whether similar biases exist across five widely used clinical speech corpora: DAIC-WoZ (depression), TORGO (dysarthria), Neurovoz (PD), MDVR-KCL (PD), and UCLASS (stuttering). For each dataset, we compare classification using the first second of audio, silent segments, and full recordings, and evaluate both raw and denoised signals. Across all datasets, silence-only classification frequently matched or exceeded full-audio performance, suggesting that classification performance may be influenced by dataset-specific confounds in addition to disorder-related speech characteristics. These findings question the reliability and generalisability of speech-based biomarkers and call for stricter methodological reporting, preprocessing transparency, and bias mitigation.

eess.AS

StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation

Detecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, where speech embeddings from overlapping windows are represented as graph nodes. We refine the connections between nodes using a pseudo-oracle classifier trained on weak (utterance-level) labels, with its influence controlled by an uncertainty measure from Monte Carlo dropout. Additionally, we extend the weakly labelled FluencyBank dataset by incorporating frame-level dysfluency boundaries for four dysfluency types. This provides a more realistic benchmark compared to synthetic datasets. Experiments on real and synthetic datasets show that StutterCut outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection.

cs.SD

Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example

Speech anonymisation aims to protect speaker identity by changing personal identifiers in speech while retaining linguistic content. Current methods fail to retain prosody and unique speech patterns found in elderly and pathological speech domains, which is essential for remote health monitoring. To address this gap, we propose a voice conversion-based method (DDSP-QbE) using differentiable digital signal processing and query-by-example. The proposed method, trained with novel losses, aids in disentangling linguistic, prosodic, and domain representations, enabling the model to adapt to uncommon speech patterns. Objective and subjective evaluations show that DDSP-QbE significantly outperforms the voice conversion state-of-the-art concerning intelligibility, prosody, and domain preservation across diverse datasets, pathologies, and speakers while maintaining quality and speaker anonymity. Experts validate domain preservation by analysing twelve clinically pertinent domain attributes.

cs.AI