SearcharxivSearch

arXiv subjects

Maximilian Schmitt

Publications and source records attributed to Maximilian Schmitt.

14 recordsLinked to original sources

Am I Blue or Is My Hobby Counting Teardrops? Expression Leakage in Large Language Models as a Symptom of Irrelevancy Disruption

Large language models (LLMs) have advanced natural language processing (NLP) skills such as through next-token prediction and self-attention, but their ability to integrate broad context also makes them prone to incorporating irrelevant information. Prior work has focused on semantic leakage, bias introduced by semantically irrelevant context. In this paper, we introduce expression leakage, a novel phenomenon where LLMs systematically generate sentimentally charged expressions that are semantically unrelated to the input context. To analyse the expression leakage, we collect a benchmark dataset along with a scheme to automatically generate a dataset from free-form text from common-crawl. In addition, we propose an automatic evaluation pipeline that correlates well with human judgment, which accelerates the benchmarking by decoupling from the need of annotation for each analysed model. Our experiments show that, as the model scales in the parameter space, the expression leakage reduces within the same LLM family. On the other hand, we demonstrate that expression leakage mitigation requires specific care during the model building process, and cannot be mitigated by prompting. In addition, our experiments indicate that, when negative sentiment is injected in the prompt, it disrupts the generation process more than the positive sentiment, causing a higher expression leakage rate.

cs.CL

Singular cohomology of symplectic quotients by circle actions and Kirwan surjectivity

Let $M$ be a symplectic manifold carrying a Hamiltonian $S^1$-action with momentum map $J:M \rightarrow \mathbb{R}$ and consider the corresponding symplectic quotient $\mathcal{M}_0:=J^{-1}(0)/S^1$. We extend Sjamaar's complex of differential forms on $\mathcal{M}_0$, whose cohomology is isomorphic to the singular cohomology $H(\mathcal{M}_0;\mathbb{R})$ of $\mathcal{M}_0$ with real coefficients, to a complex of differential forms on $\mathcal{M}_0$ associated with a partial desingularization $\widetilde{\mathcal{M}}_0$, which we call resolution differential forms. The cohomology of that complex turns out to be isomorphic to the de Rham cohomology $H(\widetilde{ \mathcal{M}}_0)$ of $\widetilde{\mathcal{M}}_0$. Based on this, we derive a long exact sequence involving both $H(\mathcal{M}_0;\mathbb{R})$ and $H(\widetilde{ \mathcal{M}}_0)$ and give conditions for its splitting. We then define a Kirwan map $\mathcal{K}:H_{S^1}(M) \rightarrow H(\widetilde{\mathcal{M}}_0)$ from the equivariant cohomology $H_{S^1}(M)$ of $M$ to $H(\widetilde{\mathcal{M}}_0)$ and show that its image contains the image of $H(\mathcal{M}_0;\mathbb{R})$ in $H(\widetilde{\mathcal{M}}_0)$ under the natural inclusion. Combining both results in the case that all fixed point components of $M$ have vanishing odd cohomology we obtain a surjection $\check κ:H^\textrm{ev}_{S^1}(M) \rightarrow H^\textrm{ev}(\mathcal{M}_0;\mathbb{R})$ in even degrees, while already simple examples show that a similar surjection in odd degrees does not exist in general. As an interesting class of examples we study abelian polygon spaces.

math.SG

Dawn of the transformer era in speech emotion recognition: closing the valence gap

Recent advances in transformer-based architectures which are pre-trained in self-supervised manner have shown great promise in several machine learning tasks. In the audio domain, such architectures have also been successfully utilised in the field of speech emotion recognition (SER). However, existing works have not evaluated the influence of model size and pre-training data on downstream performance, and have shown limited attention to generalisation, robustness, fairness, and efficiency. The present contribution conducts a thorough analysis of these aspects on several pre-trained variants of wav2vec 2.0 and HuBERT that we fine-tuned on the dimensions arousal, dominance, and valence of MSP-Podcast, while additionally using IEMOCAP and MOSI to test cross-corpus generalisation. To the best of our knowledge, we obtain the top performance for valence prediction without use of explicit linguistic information, with a concordance correlation coefficient (CCC) of .638 on MSP-Podcast. Furthermore, our investigations reveal that transformer-based architectures are more robust to small perturbations compared to a CNN-based baseline and fair with respect to biological sex groups, but not towards individual speakers. Finally, we are the first to show that their extraordinary success on valence is based on implicit linguistic information learnt during fine-tuning of the transformer layers, which explains why they perform on-par with recent multimodal approaches that explicitly utilise textual information. Our findings collectively paint the following picture: transformer-based architectures constitute the new state-of-the-art in SER, but further advances are needed to mitigate remaining robustness and individual speaker issues. To make our findings reproducible, we release the best performing model to the community.

eess.AS

Probing Speech Emotion Recognition Transformers for Linguistic Knowledge

Large, pre-trained neural networks consisting of self-attention layers (transformers) have recently achieved state-of-the-art results on several speech emotion recognition (SER) datasets. These models are typically pre-trained in self-supervised manner with the goal to improve automatic speech recognition performance -- and thus, to understand linguistic information. In this work, we investigate the extent in which this information is exploited during SER fine-tuning. Using a reproducible methodology based on open-source tools, we synthesise prosodically neutral speech utterances while varying the sentiment of the text. Valence predictions of the transformer model are very reactive to positive and negative sentiment content, as well as negations, but not to intensifiers or reducers, while none of those linguistic features impact arousal or dominance. These findings show that transformers can successfully leverage linguistic information to improve their valence predictions, and that linguistic analysis should be included in their testing.

cs.CL

On the signature of biquotients

We generalize Hirzebruch's computation of the signature of equal rank homogeneous spaces to a large class of biquotients.

math.DG

A German Corpus for Fine-Grained Named Entity Recognition and Relation Extraction of Traffic and Industry Events

Monitoring mobility- and industry-relevant events is important in areas such as personal travel planning and supply chain management, but extracting events pertaining to specific companies, transit routes and locations from heterogeneous, high-volume text streams remains a significant challenge. This work describes a corpus of German-language documents which has been annotated with fine-grained geo-entities, such as streets, stops and routes, as well as standard named entity types. It has also been annotated with a set of 15 traffic- and industry-related n-ary relations and events, such as accidents, traffic jams, acquisitions, and strikes. The corpus consists of newswire texts, Twitter messages, and traffic reports from radio stations, police and railway companies. It allows for training and evaluating both named entity recognition algorithms that aim for fine-grained typing of geo-entities, as well as n-ary relation extraction systems.

cs.CL

SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild

Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our life. Accurately annotated real-world data are the crux in devising such systems. However, existing databases usually consider controlled settings, low demographic variability, and a single task. In this paper, we introduce the SEWA database of more than 2000 minutes of audio-visual data of 398 people coming from six cultures, 50% female, and uniformly spanning the age range of 18 to 65 years old. Subjects were recorded in two different contexts: while watching adverts and while discussing adverts in a video chat. The database includes rich annotations of the recordings in terms of facial landmarks, facial action units (FAU), various vocalisations, mirroring, and continuously valued valence, arousal, liking, agreement, and prototypic examples of (dis)liking. This database aims to be an extremely valuable resource for researchers in affective computing and automatic human sensing and is expected to push forward the research in human behaviour analysis, including cultural studies. Along with the database, we provide extensive baseline experiments for automatic FAU detection and automatic valence, arousal and (dis)liking intensity estimation.

cs.HC

AVEC 2019 Workshop and Challenge: State-of-Mind, Detecting Depression with AI, and Cross-Cultural Affect Recognition

The Audio/Visual Emotion Challenge and Workshop (AVEC 2019) "State-of-Mind, Detecting Depression with AI, and Cross-cultural Affect Recognition" is the ninth competition event aimed at the comparison of multimedia processing and machine learning methods for automatic audiovisual health and emotion analysis, with all participants competing strictly under the same conditions. The goal of the Challenge is to provide a common benchmark test set for multimodal information processing and to bring together the health and emotion recognition communities, as well as the audiovisual processing communities, to compare the relative merits of various approaches to health and emotion recognition from real-life data. This paper presents the major novelties introduced this year, the challenge guidelines, the data used, and the performance of the baseline systems on the three proposed tasks: state-of-mind recognition, depression assessment with AI, and cross-cultural affect sensing, respectively.

cs.HC

Weakly Supervised One-Shot Detection with Attention Similarity Networks

Neural network models that are not conditioned on class identities were shown to facilitate knowledge transfer between classes and to be well-suited for one-shot learning tasks. Following this motivation, we further explore and establish such models and present a novel neural network architecture for the task of weakly supervised one-shot detection. Our model is only conditioned on a single exemplar of an unseen class and a larger target example that may or may not contain an instance of the same class as the exemplar. By pairing a Siamese similarity network with an attention mechanism, we design a model that manages to simultaneously identify and localise instances of classes unseen at training time. In experiments with datasets from the computer vision and audio domains, the proposed method considerably outperforms the baseline methods for the weakly supervised one-shot detection task.

stat.ML

Active Brownian motion of emulsion droplets: Coarsening dynamics at the interface and rotational diffusion

A micron-sized droplet of bromine water immersed in a surfactant-laden oil phase can swim (S. Thutupalli, R. Seemann, S. Herminghaus, New J. Phys. 13 073021 (2011)). The bromine reacts with the surfactant at the droplet interface and generates a surfactant mixture. It can spontaneously phase-separate due to solutocapillary Marangoni flow, which propels the droplet. We model the system by a diffusion-advection-reaction equation for the mixture order parameter at the interface including thermal noise and couple it to fluid flow. Going beyond previous work, we illustrate the coarsening dynamics of the surfactant mixture towards phase separation in the axisymmetric swimming state. Coarsening proceeds in two steps: an initially slow growth of domain size followed by a nearly ballistic regime. On larger time scales thermal fluctuations in the local surfactant composition initiates random changes in the swimming direction and the droplet performs a persistent random walk, as observed in experiments. Numerical solutions show that the rotational correlation time scales with the square of the inverse noise strength. We confirm this scaling by a perturbation theory for the fluctuations in the mixture order parameter and thereby identify the active emulsion droplet as an active Brownian particle.

cond-mat.soft

openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit

We introduce openXBOW, an open-source toolkit for the generation of bag-of-words (BoW) representations from multimodal input. In the BoW principle, word histograms were first used as features in document classification, but the idea was and can easily be adapted to, e.g., acoustic or visual low-level descriptors, introducing a prior step of vector quantisation. The openXBOW toolkit supports arbitrary numeric input features and text input and concatenates computed subbags to a final bag. It provides a variety of extensions and options. To our knowledge, openXBOW is the first publicly available toolkit for the generation of crossmodal bags-of-words. The capabilities of the tool are exemplified in two sample scenarios: time-continuous speech-based emotion recognition and sentiment analysis in tweets where improved results over other feature representation forms were observed.

cs.CV

Marangoni flow at droplet interfaces: Three-dimensional solution and applications

The Marangoni effect refers to fluid flow induced by a gradient in surface tension at a fluid-fluid interface. We determine the full three-dimensional Marangoni flow generated by a non-uniform surface tension profile at the interface of a self-propelled spherical emulsion droplet. For all flow fields inside, outside, and at the interface of the droplet, we give analytical formulas. We also calculate the droplet velocity vector $\mathbf{v}^D$, which describes the swimming kinematics of the droplet, and generalize the squirmer parameter $β$, which distinguishes between different swimmer types called neutral, pusher, or puller. In the second part of this paper, we present two illustrative examples, where the Marangoni effect is used in active emulsion droplets. First, we demonstrate how micelle adsorption can spontaneously break the isotropic symmetry of an initially surfactant-free emulsion droplet, which then performs directed motion. Second, we think about light-switchable surfactants and laser light to create a patch with a different surfactant type at the droplet interface. Depending on the setup such as the wavelength of the laser light and the surfactant type in the outer bulk fluid, one can either push droplets along unstable trajectories or pull them along straight or oscillatory trajectories regulated by specific parameters. We explore these cases for strongly absorbing and for transparent droplets.

cond-mat.soft

Swimming active droplet: A theoretical analysis

Recently, an active microswimmer was constructed where a micron-sized droplet of bromine water was placed into a surfactant-laden oil phase. Due to a bromination reaction of the surfactant at the interface, the surface tension locally increases and becomes non-uniform. This drives a Marangoni flow which propels the squirming droplet forward. We develop a diffusion-advection-reaction equation for the order parameter of the surfactant mixture at the droplet interface using a mixing free energy. Numerical solutions reveal a stable swimming regime above a critical Marangoni number M but also stopping and oscillating states when M is increased further. The swimming droplet is identified as a pusher whereas in the oscillating state it oscillates between being a puller and a pusher.

cond-mat.soft

Modelling bacterial flagellar growth

The growth of bacterial flagellar filaments is a self-assembly process where flagellin molecules are transported through the narrow core of the flagellum and are added at the distal end. To model this situation, we generalize a growth process based on the TASEP model by allowing particles to move both forward and backward on the lattice. The bias in the forward and backward jump rates determines the lattice tip speed, which we analyze and also compare to simulations. For positive bias, the system is in a non-equilibrium steady state and exhibits boundary-induced phase transitions. The tip speed is constant. In the no-bias case we find that the length of the lattice grows as $N(t)\propto\sqrt{t}$, whereas for negative drift $N(t)\propto\ln{t}$. The latter result agrees with experimental data of bacterial flagellar growth.

physics.bio-ph