Searcharxiv⌕ Search

arXiv subjects

Line Katrine Harder Clemmensen

Publications and source records attributed to Line Katrine Harder Clemmensen.

4 recordsLinked to original sources

Linguistic Distance Segregates Latent Representations in Automatic Speech Recognition Systems

While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers' L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models ($p<0.001$ across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures

cs.CL↗

Classification of Tennis Actions Using Deep Learning

Recent advances of deep learning makes it possible to identify specific events in videos with greater precision. This has great relevance in sports like tennis in order to e.g., automatically collect game statistics, or replay actions of specific interest for game strategy or player improvements. In this paper, we investigate the potential and the challenges of using deep learning to classify tennis actions. Three models of different size, all based on the deep learning architecture SlowFast were trained and evaluated on the academic tennis dataset THETIS. The best models achieve a generalization accuracy of 74 %, demonstrating a good performance for tennis action classification. We provide an error analysis for the best model and pinpoint directions for improvement of tennis datasets in general. We discuss the limitations of the data set, general limitations of current publicly available tennis data-sets, and future steps needed to make progress.

cs.CV↗

Applying Pre-Trained Deep-Learning Model on Wrist Angel Data -- An Analysis Plan

We aim to investigate if we can improve predictions of stress caused by OCD symptoms using pre-trained models, and present our statistical analysis plan in this paper. With the methods presented in this plan, we aim to avoid bias from data knowledge and thereby strengthen our hypotheses and findings. The Wrist Angel study, which this statistical analysis plan concerns, contains data from nine participants, between 8 and 17 years old, diagnosed with obsessive-compulsive disorder (OCD). The data was obtained by an Empatica E4 wristband, which the participants wore during waking hours for 8 weeks. The purpose of the study is to assess the feasibility of predicting the in-the-wild OCD events captured during this period. In our analysis, we aim to investigate if we can improve predictions of stress caused by OCD symptoms, and to do this we have created a pre-trained model, trained on four open-source data for stress prediction. We intend to apply this pre-trained model to the Wrist Angel data by fine-tuning, thereby utilizing transfer learning. The pre-trained model is a convolutional neural network that uses blood volume pulse, heart rate, electrodermal activity, and skin temperature as time series windows to predict OCD events. Furthermore, using accelerometer data, another model filters physical activity to further improve performance, given that physical activity is physiologically similar to stress. By evaluating various ways of applying our model (fine-tuned, non-fine-tuned, pre-trained, non-pre-trained, and with or without activity classification), we contextualize the problem such that it can be assessed if transfer learning is a viable strategy in this domain.

stat.AP↗

Pre-processing Blood-Volume-Pulse for In-the-wild Applications

Blood-volume-pulse (BVP) is a biosignal commonly used in applications for non-invasive affect recognition and wearable technology. However, its predisposition to noise constitutes limitations for its application in real-life settings. This paper revisits BVP processing and proposes standard practices for feature extraction from empirical observations of BVP. We propose a method for improving the use of features in the presence of noise and compare it to a standard signal processing approach of a 4th order Butterworth bandpass filter with cut-off frequencies of 1 Hz and 8 Hz. Our method achieves better results for most time features as well as for a subset of the frequency features. We find that all but one time feature and around half of the frequency features perform better when the noisy parts are known (best case). When the noisy parts are unknown and estimated using a metric of skewness, the proposed method in general works better or similar to the Butterworth bandpass filter, but both methods also fail for a subset features. Our results can be used to select BVP features that are meaningful under different SNR conditions.

eess.SP↗