SearcharxivSearch

arXiv subjects

Roksana Khanom

Publications and source records attributed to Roksana Khanom.

2 recordsLinked to original sources

A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech

Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by language-specific phonetic variation. We present CL-DAF, a Cross-Lingual Disease-Alignment Framework that identifies acoustic dimensions whose disease effects remain consistent across languages. Using 201 English and 75 newly collected Bangla speakers, we construct a common 272-dimensional acoustic representation and quantify disease alignment using signed rank-biserial effects and the Language Invariance Score. We first show that spontaneous Bangla speech separates COPD from controls (AUC 0.85); however, 133 features reverse their disease direction across languages and the full representation transfers poorly (AUC 0.49 from Bangla to English). CL-DAF isolates 26 disease-aligned features that raise AUCs to 0.825 and 0.722 from English to Bangla and Bangla to English, respectively. These findings provide a foundation for multilingual clinical speech models emphasizing pathology over language-dependent variation.

cs.SD

SpiroPhonia: Non-Invasive Respiratory Health Assessment from Spontaneous Speech

Chronic Obstructive Pulmonary Disease (COPD) remains a major global health challenge, emphasizing the need for accessible and non-invasive detection. Since speech production is fundamentally linked to respiratory physiology, its disruptions can serve as indirect indicators of pulmonary impairment. This study introduces SpiroPhonia, a machine learning framework that leverages spontaneous speech for respiratory health assessment. We evaluated SpiroPhonia on a new dataset of 201 speakers (102 with COPD, 99 healthy controls). By integrating statistical analysis with recursive feature selection, we identified a compact set of discriminative speech markers. Our best model achieved 78% accuracy, 80% F1-score, and 87% AUC. This performance on spontaneous speech is competitive with methods using controlled laboratory recordings. Findings demonstrate that everyday speech encodes robust respiratory biomarkers, paving the way for continuous health monitoring via voice-enabled technologies.

cs.SD