SearcharxivSearch

arXiv subjects

Jens Madsen

Publications and source records attributed to Jens Madsen.

11 recordsLinked to original sources

Towards Quantifying Benchmark Optimization in ASR Models

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

cs.SD

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

cs.SD

From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition

The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analysis, understanding of human affect and behavior in real-world, unconstrained environments. The workshop maintains its dual structure, comprising both a competition and a paper track. The ABAW Competition introduces a diverse set of challenges targeting key aspects of affective and behavioral understanding, including continuous affect (valence-arousal) estimation, discrete affect (expression and action unit) recognition, as well as more complex behavior analysis tasks, such as emotional mimicry intensity estimation, ambivalence/hesitancy recognition and fine-grained violence detection. These challenges are built upon large-scale in-the-wild datasets, providing comprehensive benchmarks for state-of-the-art approaches. In parallel, the paper track presents a wide range of contributions spanning pose, motion & behavior estimation, affect modelling & multimodal learning, benchmarks, datasets & evaluation protocols, fairness, robustness & deployment. Overall, the 10th ABAW Workshop and Competition continues to serve as a key platform for benchmarking, collaboration and innovation, shaping the development of next-generation multimodal, human-centered AI systems.

cs.CV

The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge

The 2026 ACII Dyadic Conversations (ACII-DaiKon) Workshop & Challenge introduces a benchmark for modeling interpersonal affect and social dynamics in dyadic conversations. Although conversational affect modeling has advanced rapidly, most benchmarks remain speaker-centric and underrepresent coupled, time-evolving processes between partners, including directional influence, conversational timing coordination, and rapport development. To address this gap, ACII-DaiKon presents three coordinated sub-challenges built on a shared dataset: (1) directional interpersonal influence prediction, (2) turn-taking prediction (next-speaker and time-to-next-speech), and (3) rapport trajectory prediction across full interactions. The challenge is built on the Hume-DaiKon dataset, comprising 945 dyadic conversations (743.4 hours of audiovisual data) collected under naturalistic conditions across five languages. The benchmark supports multimodal modeling, temporal reasoning, and cross-context generalization through fixed train/validation/test splits, standardized metrics, and released baseline systems. Evaluation uses Concordance Correlation Coefficient (CCC), Pearson correlation, Macro-F1, and Mean Absolute Error (MAE) depending on the sub-challenge. Baseline experiments establish initial reference performance, with best test results of 0.40 CCC and 0.50 Pearson for influence prediction, 0.66 Macro-F1 and 1.50~s MAE for turn-taking, and 0.68 CCC and 0.70 Pearson for rapport trajectory modeling. These results indicate that while current methods capture coarse dyadic patterns, robust modeling of directional dependence and long-horizon interpersonal dynamics remains challenging. The workshop provides a shared platform for rigorous comparison and cross-disciplinary discussion on data validity, evaluation protocols, and culturally aware modeling for dyadic interaction.

cs.AI

Real-time estimation of overt attention from dynamic features of the face using deep-learning

Students often drift in and out of focus during class. Effective teachers recognize this and re-engage them when necessary. With the shift to remote learning, teachers have lost the visual feedback needed to adapt to varying student engagement. We propose using readily available front-facing video to infer attention levels based on movements of the eyes, head, and face. We train a deep learning model to predict a measure of attention based on overt eye movements. Specifically, we measure Inter-Subject Correlation of eye movements in ten-second intervals while students watch the same educational videos. In 3 different experiments (N=83) we show that the trained model predicts this objective metric of attention on unseen data with $R^2$=0.38, and on unseen subjects with $R^2$=0.26-0.30. The deep network relies mostly on a student's eye movements, but to some extent also on movements of the brows, cheeks, and head. In contrast to Inter-Subject Correlation of the eyes, the model can estimate attentional engagement from individual students' movements without needing reference data from an attentive group. This enables a much broader set of online applications. The solution is lightweight and can operate on the client side, which mitigates some of the privacy concerns associated with online attention monitoring. GitHub implementation is available at https://github.com/asortubay/timeISC

cs.CV

VARX Granger Analysis: Modeling, Inference, and Applications

Complex systems, such as brains, markets, and societies, exhibit internal dynamics influenced by external factors. Disentangling delayed external effects from internal dynamics within these systems is often challenging. We propose using a Vector Autoregressive model with eXogenous input (VARX) to capture delayed interactions between internal and external variables. While this model aligns with Granger's statistical formalism for testing "causal relations", the connection between the two is not widely understood. Here, we bridge this gap by providing fundamental equations, user-friendly code, and demonstrations using simulated and real-world data from neuroscience, physiology, sociology, and economics. Our examples illustrate how the model avoids spurious correlation by factoring out external influences from internal dynamics, leading to more parsimonious explanations of the systems. We also provide methods for enhancing model efficiency, such as L2 regularization for limited data and basis functions to cope with extended delays. Additionally, we analyze model performance under various scenarios where model assumptions are violated. MATLAB, Python, and R code are provided for easy adoption: https://github.com/lcparra/varx

stat.ME

Predicting Changes in Affective States using Neural Networks

Knowledge of patients affective state could prove to be crucial for health-care professionals in both diagnosis and treatment, however, this requires patients to report how they feel. In practice the sampling rate of affective states needs to be kept low, in order to ensure that the patients can rest. Furthermore using traditional methods of measuring affective states, is not always possible, e.g. patients can be incapable of verbal communications. In this study we explore the prediction of peoples self-reported affective state by measuring multiple physiological signals. We use different Neural networks (NN) setups and compare with different multiple linear regression (MLR) setups for prediction of changes in affective states. The results showed that NN and MLR predicted the change in affective states with accuracies of 91.88% and 89.10%, respectively.

cs.CY

Collisional transport across the magnetic field in drift-fluid models

Drift ordered fluid models are widely applied in studies of low-frequency turbulence in the edge and scrape-off layer regions of magnetically confined plasmas. Here, we show how collisional transport across the magnetic field is self-consistently incorporated into drift-fluid models without altering the drift-fluid energy integral. We demonstrate that the inclusion of collisional transport in drift-fluid models gives rise to diffusion of particle density, momentum and pressures in drift-fluid turbulence models and thereby obviate the customary use of artificial diffusion in turbulence simulations. We further derive a computationally efficient, two-dimensional model which can be time integrated for several turbulence de-correlation times using only limited computational resources. The model describes interchange turbulence in a two-dimensional plane perpendicular to the magnetic field located at the outboard midplane of a tokamak. The model domain has two regions modeling open and closed field lines. The model employs a computational expedient model for collisional transport. Numerical simulations show good agreement between the full and the simplified model for collisional transport.

physics.plasm-ph

Verification of BOUT++ by the Method of Manufactured Solutions

BOUT++ is a software package designed for solving plasma fluid models. It has been used to simulate a wide range of plasma phenomena ranging from linear stability analysis to 3D plasma turbulence, and is capable of simulating a wide range of drift-reduced plasma fluid and gyro-fluid models. A verification exercise has been performed as part of a EUROfusion Enabling Research project, to rigorously test the correctness of the algorithms implemented in BOUT++, by testing order-of-accuracy convergence rates using the Method of Manufactured Solutions (MMS). We present tests of individual components including time-integration and advection schemes, non-orthogonal coordinate systems and the shifted metric procedure which is used to handle highly sheared grids. The Flux Coordinate Independent (FCI) approach to differencing along magnetic field-lines has been implemented in BOUT++, and is here verified using the MMS in a sheared slab configuration. Finally we show tests of three complete models: 2-field Hasegawa-Wakatani, 3-field reduced MHD in 3D toroidal coordinates, and 5-field reduced MHD in slab geometry.

physics.plasm-ph

Measurement of a 2D fast-ion velocity distribution function by tomographic inversion of fast-ion D-alpha spectra

We present the first measurement of a local fast-ion 2D velocity distribution function $f(v_\parallel, v_\perp)$. To this end, we heated a plasma in ASDEX Upgrade by neutral beam injection and measured spectra of fast-ion D-alpha (FIDA) light from the plasma center in three views simultaneously. The measured spectra agree very well with synthetic spectra calculated from a TRANSP/NUBEAM simulation. Based on the measured FIDA spectra alone, we infer $f(v_\parallel, v_\perp)$ by tomographic inversion. Salient features of our measurement of $f(v_\parallel, v_\perp)$ agree reasonably well with the simulation: the measured as well as the simulated $f(v_\parallel, v_\perp)$ are lopsided towards negative velocities parallel to the magnetic field, and they have similar shapes. Further, the peaks in the simulation of $f(v_\parallel, v_\perp)$ at full and half injection energies of the neutral beam also appear in the measurement at similar velocity-space locations. We expect that we can measure spectra in up to seven views simultaneously in the next ASDEX Upgrade campaign which would further improve measurements of $f(v_\parallel, v_\perp)$ by tomographic inversion.

physics.plasm-ph

Doppler tomography in fusion plasmas and astrophysics

Doppler tomography is a well-known method in astrophysics to image the accretion flow, often in the shape of thin discs, in compact binary stars. As accretion discs rotate, all emitted line radiation is Doppler-shifted. In fast-ion D-alpha (FIDA) spectroscopy measurements in magnetically confined plasma, the D-alpha-photons are likewise Doppler-shifted ultimately due to gyration of the fast ions. In either case, spectra of Doppler-shifted line emission are sensitive to the velocity distribution of the emitters. Astrophysical Doppler tomography has lead to images of accretion discs of binaries revealing bright spots, spiral structures, and flow patterns. Fusion plasma Doppler tomography has lead to an image of the fast-ion velocity distribution function in the tokamak ASDEX Upgrade. This image matched numerical simulations very well. Here we discuss achievements of the Doppler tomography approach, its promise and limits, analogies and differences in astrophysical and fusion plasma Doppler tomography, and what can be learned by comparison of these applications.

astro-ph.SR