SearcharxivSearch

arXiv subjects

Bernhard U. Seeber

Publications and source records attributed to Bernhard U. Seeber.

8 recordsLinked to original sources

Head, posture, and full-body gestures in unscripted dyadic conversations in noise

Visual prosody may be critical for communication success in face-to-face conversations in noisy settings. Here, we explore the involvement of hand, head, and whole-body movements, as well as gesturing quality, in dyadic conversations in noisy settings. We hypothesize that increasing background noise would alter the frequency of conversation-related movements to support the roles of the speaker and the listener. Specifically, talkers may increase gesticulation and thus the use of hand, head, trunk, or leg movements more often, while listeners may increase backchanneling or head and trunk movements to improve the signal-to-noise ratio. Additionally, we test whether the synchrony between speech and hand gestures is affected by background noise. Here, pairs of normal hearing participants (n=8) stood in an audiovisual virtual environment while talking freely. The conversational movements were described using a newly developed labeling system with categories that respect their communicative function. The results showed higher gesturing rate during speaking than during listening. Increased levels of background noise led to increased hand-gesture complexity, modulation of head movements, and a change in trunk movements. People spoke 0.7 dB - 1.4 dB louder during hand gesturing in comparison to times with static drop posture but this was unrelated to presence of background noise. The analysis of hand-speech synchrony showed a modest decrease in synchrony for moderate noise level. People adapt their communicative behavior to increased background noise levels by increases in speech production levels and gesturing which may drive additional increase in speech production due to biomechanical coupling; listeners may increase backchanneling to support the exchange and their own signal-to-noise ratio. The synchrony analysis may reflect motivational factors of communication in noisy environments.

cs.HC

Perceptual evaluation of Acoustic Level of Detail in Virtual Acoustic Environments

Virtual acoustic environments enable the creation and simulation of realistic and eco-logically valid daily-life situations vital for hearing research and audiology. Reverberant indoor environments are particularly important. For real-time applications, room acous-tics simulation requires simplifications, however, the necessary acoustic level of detail (ALOD) remains unclear in order to capture all perceptually relevant effects. This study examines the impact of varying ALOD in simulations of three real environments: a living room with a coupled kitchen, a pub, and an underground station. ALOD was varied by generating different numbers of image sources for early reflections, or by excluding geo-metrical room details specific for each environment. Simulations were perceptually eval-uated using headphones in comparison to binaural room impulse responses measured with a dummy head in the corresponding real environments, or by using loudspeakers. The study assessed the perceived overall difference for a pulse stimulus, a played electric bass and a speech token. Additionally, plausibility, speech intelligibility, and externaliza-tion were evaluated. Results indicate that a strong reduction in ALOD is feasible while maintaining similar plausibility, speech intelligibility, and externalization as with dummy head recordings. The number and accuracy of early reflections appear less relevant, pro-vided diffuse late reverberation is appropriately represented.

eess.AS

Tone generation in an open-end organ pipe: How a resonating sphere of air stops the pipe

According to the classical Helmholtz picture, an organ pipe while generating its eigentone has two anti-nodes at the two open ends of a cylinder, the anti-nodes being taken as boundary condition for the corresponding sound. Since 1860 it is also known that according to the classical picture the pipe actually sounds lower, which is to say that the pipe so-to-speak sounds longer than it is, a long-standing enigma. As for the pipe's end, we have resolved this acoustic enigma by detailing the physics of the airflow at the pipe's open end and showing that the boundary condition is actually the pipe's acoustically resonating vortical sphere (PARVS). The PARVS geometry entails a sound-radiating hemisphere based on the pipe's open end and enclosing a vortex ring. In this way we obtain not only a physical explanation of sound radiation from the organ-pipe's open end, in particular, of its puzzling dependence upon the pipe's radius, but also an appreciation of it as realization of the sound of the flute, mankind's oldest musical instrument.

physics.flu-dyn

Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments

In daily life, social interaction and acoustic communication often take place in complex acoustic environments (CAE) with a variety of interfering sounds and reverberation. For hearing research and the evaluation of hearing systems, simulated CAEs using virtual reality techniques have gained interest in the context of ecological validity. In the current study, the effect of scene complexity and visual representation of the scene on psychoacoustic measures like sound source location, distance perception, loudness, speech intelligibility, and listening effort in a virtual audio-visual environment was investigated. A 3-dimensional, 86-channel loudspeaker array was used to render the sound field in combination with or without a head-mounted display (HMD) to create an immersive stereoscopic visual representation of the scene. The scene consisted of a ring of eight (virtual) loudspeakers which played a target speech stimulus and nonsense speech interferers in several spatial conditions. Either an anechoic (snowy outdoor scenery) or echoic environment (loft apartment) with a reverberation time (T60) of about 1.5 s was simulated. In addition to varying the number of interferers, scene complexity was varied by assessing the psychoacoustic measures in isolated consecutive measurements orcsimultaneously. Results showed no significant effect of wearing the HMD on the data. Loudness and distance perception showed significantly different results when they were measured simultaneously instead of consecutively in isolation. The advantage of the suggested setup is that it can be directly transferred to a corresponding real room, enabling a 1:1 comparison and verification of the perception experiments in the real and virtual environment.

eess.AS

Communication conditions in virtual acoustic scenes in an underground station

Underground stations are a common communication situation in towns: we talk with friends or colleagues, listen to announcements or shop for titbits while background noise and reverberation are challenging communication. Here, we perform an acoustical analysis of two communication scenes in an underground station in Munich and test speech intelligibility. The acoustical conditions were measured in the station and are compared to simulations in the real-time Simulated Open Field Environment (rtSOFE). We compare binaural room impulse responses measured with an artificial head in the station to modeled impulse responses for free-field auralization via 60 loudspeakers in the rtSOFE. We used the image source method to model early reflections and a set of multi-microphone recordings to model late reverberation. The first communication scene consists of 12 equidistant (1.6 m) horizontally spaced source positions around a listener, simulating different direction-dependent spatial unmasking conditions. The second scene mimics an approaching speaker across six radially spaced source positions (from 1 m to 10 m) with varying direct sound level and thus direct-to-reverberant energy. The acoustic parameters of the underground station show a moderate amount of reverberation (T30 in octave bands was between 2.3 s and 0.6 s and early-decay times between 1.46 s and 0.46 s). The binaural and energetic parameters of the auralization were in a close match to the measurement. Measured speech reception thresholds were within the error of the speech test, letting us to conclude that the auralized simulation reproduces acoustic and perceptually relevant parameters for speech intelligibility with high accuracy.

cs.SD

Auditory-visual scenes for hearing research

While experimentation with synthetic stimuli in abstracted listening situations has a long standing and successful history in hearing research, an increased interest exists on closing the remaining gap towards real-life listening by replicating situations with high ecological validity in the lab. This is important for understanding the underlying auditory mechanisms and their relevance in real-life situations as well as for developing and evaluating increasingly sophisticated algorithms for hearing assistance. A range of 'classical' stimuli and paradigms have evolved to de-facto standards in psychoacoustics, which are simplistic and can be easily reproduced across laboratories. While they ideally allow for across laboratory comparisons and reproducible research, they, however, lack the acoustic stimulus complexity and the availability of visual information as observed in everyday life communication and listening situations. This contribution aims to provide and establish an extendable set of complex auditory-visual scenes for hearing research that allow for ecologically valid testing in realistic scenes while also supporting reproducibility and comparability of scientific results. Three virtual environments are provided (underground station, pub, living room), consisting of a detailed visual model, an acoustic geometry model with acoustic surface properties as well as a set of acoustic measurements in the respective real-world environments. The current data set enables i) audio-visual research in a reproducible set of environments, ii) comparison of room acoustic simulation methods with "ground truth" acoustic measurements, iii) a condensation point for future extensions and contributions for developments towards standardized test cases for ecologically valid hearing research in complex scenes.

physics.med-ph

Fast processing explains the effect of sound reflection on binaural unmasking

Sound reflections and late reverberation alter energetic and binaural cues of a target source, thereby affecting it's detection in noise. Two experiments investigated detection of harmonic complex tones, centered around 500 Hz, in noise in a virtual room with different modifications of simulated room impulse responses (RIR). Stimuli were auralized using the SOFE's loudspeakers in anechoic space. The target was presented from the front or at 0$^\circ$ azimuth, while an anechoic noise masker was simultaneously presented at 0$^\circ$. In the first experiment, early reflections were progressively added to the RIR and detection thresholds of the reverberant target were measured. For a frontal sound source, detection thresholds decreased while adding the first 45 ms of early reflections, whereas for a lateral sound source thresholds remained constant. In the second experiment, early reflections were cut out while late reflections were kept along with the direct sound. Results for a target at 0$^\circ$ show that even reflections as late as 150 ms reduce detection thresholds compared to only the direct sound. A binaural model with a sluggishness component following the computation of binaural unmasking in short windows predicts measured and literature results better than when large windows are used.

eess.AS

MP3 Compression To Diminish Adversarial Noise in End-to-End Speech Recognition

Audio Adversarial Examples (AAE) represent specially created inputs meant to trick Automatic Speech Recognition (ASR) systems into misclassification. The present work proposes MP3 compression as a means to decrease the impact of Adversarial Noise (AN) in audio samples transcribed by ASR systems. To this end, we generated AAEs with the Fast Gradient Sign Method for an end-to-end, hybrid CTC-attention ASR system. Our method is then validated by two objective indicators: (1) Character Error Rates (CER) that measure the speech decoding performance of four ASR models trained on uncompressed, as well as MP3-compressed data sets and (2) Signal-to-Noise Ratio (SNR) estimated for both uncompressed and MP3-compressed AAEs that are reconstructed in the time domain by feature inversion. We found that MP3 compression applied to AAEs indeed reduces the CER when compared to uncompressed AAEs. Moreover, feature-inverted (reconstructed) AAEs had significantly higher SNRs after MP3 compression, indicating that AN was reduced. In contrast to AN, MP3 compression applied to utterances augmented with regular noise resulted in more transcription errors, giving further evidence that MP3 encoding is effective in diminishing only AN.

eess.AS