SearcharxivSearch

arXiv subjects

Jinkook Lee

Publications and source records attributed to Jinkook Lee.

3 recordsLinked to original sources

Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults

Automatic language proficiency assessment in the context of multilingual interview-based settings remains underexplored. In this work, we develop Whisper-based speaker-role and language diarization systems to automatically extract respondent speech and characterize language usage in multilingual interviews with older adults. We further investigate whether diarization-derived conversational and language-use behaviors can support downstream language proficiency assessment. Results show that language-adapted Whisper models substantially improve language diarization performance for lower-resource and linguistically related Indian languages. Statistical analyses reveal that respondent speech ratio and intended language usage are strong predictors of proficiency ratings. Furthermore, simple diarization-derived behavioral features achieve performance comparable to Whisper-based speech embeddings for proficiency prediction, while combining both yields the best results. Importantly, both the speech and language use statistical analyses and language proficiency prediction performance remain largely preserved when using fully automatic diarization outputs, demonstrating the potential of respondent-centric conversational analysis for scalable language proficiency assessment.

eess.AS

The Effects of Higher Education on Midlife Depression: Quasi-Experimental Evidence from South Korea

Higher education has expanded worldwide, with women outpacing men in many regions. While educational attainment is consistently linked to better physical health, its mental health effects - particularly for women - remain underexplored, and causal evidence is limited. We estimate the impact of college completion on depression among middle-aged women in South Korea, leveraging the 1993 higher education reform, which raised women's college attainment by 45 percentage points (pp) over the following decade. We use two nationally representative datasets to triangulate evidence, including the Korea National Health and Nutrition Examination Survey (KNHANES, 2007-2021) for physician-diagnosed depression, and the Korean Longitudinal Survey of Women and Families (KLoWF, 2007-2022) to validate findings using self-reports of depressive symptoms. We implement two-stage least squares (2SLS) with a birth-cohort instrument based on exposure to the reform (within 3 years of the cutoff in KNHANES and within 1 to 3 years in KLoWF). In KNHANES, college completion lowers physician-diagnosed depression by 2.4 pp, attenuating to 1.6 pp after adjusting for income, employment, and physical health. In KLoWF, college completion improves self-reported mental health. The weekly depressive-symptoms composite declines by 17.4 pp, attenuating to 16.4 pp after covariate adjustment. Placebo tests on unaffected cohorts yield null results. This study contributes to the growing quasi-experimental literature on education and mental health with convergent evidence across clinical diagnoses and self-reported depressive symptoms in South Korea. By focusing on college education in a non-Western setting, it extends the external validity of existing findings and highlights educational policy as a potential lever to reduce the burden of midlife depression among women.

econ.GN

VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge

In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals with Alzheimer's Disease (AD) or mild cognitive impairment (MCI) often produce measurable speech characteristics, such as slower articulation rate and lengthened sounds, in a manner similar to dialectal phonetic variations seen in speech. Building on this idea, we introduce VoxCog, an end-to-end framework that uses pre-trained dialect models to detect AD or MCI without relying on additional modalities such as text or images. Through experiments on multiple multilingual datasets for AD and MCI detection, we demonstrate that model initialization with a dialect classifier on top of speech foundation models consistently improves the predictive performance of AD or MCI. Our trained models yield similar or often better performance compared to previous approaches that ensembled several computational methods using different signal modalities. Particularly, our end-to-end speech-based model achieves 87.5% and 85.9% accuracy on the ADReSS 2020 challenge and ADReSSo 2021 challenge test sets, outperforming existing solutions that use multimodal ensemble-based computation or LLMs.

cs.SD