SearcharxivSearch

arXiv subjects

Joachim Wagner

Publications and source records attributed to Joachim Wagner.

14 recordsLinked to original sources

SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection

We introduce SynBullying, a synthetic multi-LLM conversational dataset for studying and detecting cyberbullying (CB). SynBullying provides a scalable and ethically safe alternative to human data collection by leveraging large language models (LLMs) to simulate realistic bullying interactions. The dataset offers (i) conversational structure, capturing multi-turn exchanges rather than isolated posts; (ii) context-aware annotations, where harmfulness is assessed within the conversational flow considering context, intent, and discourse dynamics; and (iii) fine-grained labeling, covering various CB categories for detailed linguistic and behavioral analysis. We evaluate SynBullying across five dimensions, including conversational structure, lexical patterns, sentiment/toxicity, role dynamics, harm intensity, and CB-type distribution. We further examine its utility by testing its performance as standalone training data and as an augmentation source for CB classification.

cs.AI

Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection

Cyberbullying (CB) presents a pressing threat, especially to children, underscoring the urgent need for robust detection systems to ensure online safety. While large-scale datasets on online abuse exist, there remains a significant gap in labeled data that specifically reflects the language and communication styles used by children. The acquisition of such data from vulnerable populations, such as children, is challenging due to ethical, legal and technical barriers. Moreover, the creation of these datasets relies heavily on human annotation, which not only strains resources but also raises significant concerns due to annotators exposure to harmful content. In this paper, we address these challenges by leveraging Large Language Models (LLMs) to generate synthetic data and labels. Our experiments demonstrate that synthetic data enables BERT-based CB classifiers to achieve performance close to that of those trained on fully authentic datasets (75.8% vs. 81.5% accuracy). Additionally, LLMs can effectively label authentic yet unlabeled data, allowing BERT classifiers to attain a comparable performance level (79.1% vs. 81.5% accuracy). These results highlight the potential of LLMs as a scalable, ethical, and cost-effective solution for generating data for CB detection.

cs.CL

Coupled dynamics in binary mixtures of colloidal Yukawa systems

The dynamical behavior of binary mixtures consisting of highly charged colloidal particles is studied by means of Brownian dynamics simulations. We investigate differently sized, but identically charged particles with nearly identical interactions between all species in highly dilute suspensions. Different short-time self-diffusion coefficients induce, mediated by electrostatic interactions, a coupling of both self and collective dynamics of differently sized particles: The long-time self-diffusion coefficients of a larger species are increased by the presence of a more mobile, smaller species and vice versa. Similar coupling effects are observed in collective dynamics where in addition to the time constant of intermediate scattering function's initial decay its functional form, quantified by exponents of a stretched exponential decay, are influenced by the presence of a differently sized species. We provide a systematic analysis of coupling effects in dependence on the ratio of sizes, number densities, and the strength of electrostatic interactions.

cond-mat.soft

Geometric measures of uniaxial solids of revolution in ${\mathbb{R}^{4}}$ and their relation to the second virial coefficient

We provide analytical expressions for the second virial coefficients of hard, convex, monoaxial solids of revolution in ${\mathbb{R}^{4}}$. The excluded volume per particle and thus the second virial coefficient is calculated using quermassintegrals and rotationally invariant mixed volumes based on the Brunn-Minkowski theorem. We derive analytical expressions for the mutual excluded volume of four-dimensional hard solids of revolution in dependence on their aspect ratio $\nu$ including the limits of infinitely thin oblate and infinitely long prolate geometries. Using reduced second virial coefficients $B_2^{\ast}=B_2/V_{\mathrm{P}}$ as size-independent quantities with $V_{\mathrm{P}}$ denoting the $D$-dimensional particle volume, the influence of the particle geometry to the mutual excluded volume is analyzed for various shapes. Beyond the aspect ratio $\nu$, the detailed particle shape influences the reduced second virial coefficients $B_2^{\ast}$. We prove that for $D$-dimensional spherocylinders in arbitrary-dimensional Euclidean spaces ${\mathbb{R}^{D}}$ their excluded volume solely depends on at most three intrinsic volumes, whereas for different convex geometries $D$ intrinsic volumes are required. For $D$-dimensional ellipsoids of revolution, the general parity $B_2^{\ast}(\nu)=B_2^{\ast}(\nu^{-1})$ is proven.

cond-mat.stat-mech

Measurable Structure Factors of Dense Dispersions Containing Polydisperse, Optically Inhomogeneous Particles

We exemplarily investigate how optical properties of single scatterers in interacting multi-particle systems influence measurable structure factors. Both particles with linear gradients of their scattering length density and core-shell structures evoke characteristic deviations between the weighted sum $\langle S(Q)\rangle$ of partial structure factors in a multicomponent system and experimentally accessible, measurable structure factors $S_{\mathrm{M}}(Q)$. While $\langle S(Q)\rangle$ contains only structural information of self-organising systems, $S_{\mathrm{M}}(Q)$ additionally is influenced by optical properties of their constituents resulting in features such as changing amplitudes, additional peaks in the low wavevector region or splitting of higher-order maxima which are not related to structural reasons. Hence, a careful data analysis regarding size-distribution and optical properties of single scatters is mandatory to avoid a misinterpretation of measurable structure factors.

cond-mat.soft

Rescaled Mode-Coupling Scheme for the Quantitative Description of Experimentally Observed Colloid Dynamics

We describe experimentally observed collective dynamics in colloidal suspensions of model hard-sphere particles using a modified mode coupling theory (MCT). This rescaled MCT is capable to describe quantitatively the wave-vector and time-dependent diffusion in these systems. Intermediate scattering functions of liquid-like structured dispersions are determined by means of static and dynamic light scattering experiments. The structure and short-time dynamics of the systems can be described quantitatively employing a multi-component Percus-Yevick ansatz for the partial structure factors and an effective, one-component description of hydrodynamic interactions based on the semi-analytical $\delta\gamma$-expansion. Combined with a recently proposed empirical modification of MCT in which memory functions are calculated using effective structure factors at rescaled number densities, the scheme is able to model the collective dynamics over the entire accessible time and wave-vector range and predicts the volume-fraction-dependence of long-time self-diffusion coefficients and the zero-shear viscosity quantitatively. This highlights the potential of MCT as a practical tool for the quantitative analysis and prediction of experimental observations.

cond-mat.soft

Revisiting Tri-training of Dependency Parsers

We compare two orthogonal semi-supervised learning techniques, namely tri-training and pretrained word embeddings, in the task of dependency parsing. We explore language-specific FastText and ELMo embeddings and multilingual BERT embeddings. We focus on a low resource scenario as semi-supervised learning can be expected to have the most impact here. Based on treebank size and available ELMo models, we select Hungarian, Uyghur (a zero-shot language for mBERT) and Vietnamese. Furthermore, we include English in a simulated low-resource setting. We find that pretrained word embeddings make more effective use of unlabelled data than tri-training but that the two approaches can be successfully combined.

cs.CL

gaBERT -- an Irish Language Model

The BERT family of neural language models have become highly popular due to their ability to provide sequences of text with rich context-sensitive token encodings which are able to generalise well to many NLP tasks. We introduce gaBERT, a monolingual BERT model for the Irish language. We compare our gaBERT model to multilingual BERT and the monolingual Irish WikiBERT, and we show that gaBERT provides better representations for a downstream parsing task. We also show how different filtering criteria, vocabulary size and the choice of subword tokenisation model affect downstream performance. We compare the results of fine-tuning a gaBERT model with an mBERT model for the task of identifying verbal multiword expressions, and show that the fine-tuned gaBERT model also performs better at this task. We release gaBERT and related code to the community.

cs.CL

The DCU-EPFL Enhanced Dependency Parser at the IWPT 2021 Shared Task

We describe the DCU-EPFL submission to the IWPT 2021 Shared Task on Parsing into Enhanced Universal Dependencies. The task involves parsing Enhanced UD graphs, which are an extension of the basic dependency trees designed to be more facilitative towards representing semantic structure. Evaluation is carried out on 29 treebanks in 17 languages and participants are required to parse the data from each language starting from raw strings. Our approach uses the Stanza pipeline to preprocess the text files, XLMRoBERTa to obtain contextualized token representations, and an edge-scoring and labeling model to predict the enhanced graph. Finally, we run a post-processing script to ensure all of our outputs are valid Enhanced UD graphs. Our system places 6th out of 9 participants with a coarse Enhanced Labeled Attachment Score (ELAS) of 83.57. We carry out additional post-deadline experiments which include using Trankit for pre-processing, XLM-RoBERTa-LARGE, treebank concatenation, and multitask learning between a basic and an enhanced dependency parser. All of these modifications improve our initial score and our final system has a coarse ELAS of 88.04.

cs.CL

The ADAPT Enhanced Dependency Parser at the IWPT 2020 Shared Task

We describe the ADAPT system for the 2020 IWPT Shared Task on parsing enhanced Universal Dependencies in 17 languages. We implement a pipeline approach using UDPipe and UDPipe-future to provide initial levels of annotation. The enhanced dependency graph is either produced by a graph-based semantic dependency parser or is built from the basic tree using a small set of heuristics. Our results show that, for the majority of languages, a semantic dependency parser can be successfully applied to the task of parsing enhanced dependencies. Unfortunately, we did not ensure a connected graph as part of our pipeline approach and our competition submission relied on a last-minute fix to pass the validation script which harmed our official evaluation scores significantly. Our submission ranked eighth in the official evaluation with a macro-averaged coarse ELAS F1 of 67.23 and a treebank average of 67.49. We later implemented our own graph-connecting fix which resulted in a score of 79.53 (language average) or 79.76 (treebank average), which would have placed fourth in the competition evaluation.

cs.CL

Treebank Embedding Vectors for Out-of-domain Dependency Parsing

A recent advance in monolingual dependency parsing is the idea of a treebank embedding vector, which allows all treebanks for a particular language to be used as training data while at the same time allowing the model to prefer training data from one treebank over others and to select the preferred treebank at test time. We build on this idea by 1) introducing a method to predict a treebank vector for sentences that do not come from a treebank used in training, and 2) exploring what happens when we move away from predefined treebank embedding vectors during test time and instead devise tailored interpolations. We show that 1) there are interpolated vectors that are superior to the predefined ones, and 2) treebank vectors can be predicted with sufficient accuracy, for nine out of ten test languages, to match the performance of an oracle approach that knows the most suitable predefined treebank embedding for the test set.

cs.CL

Cross-lingual Parsing with Polyglot Training and Multi-treebank Learning: A Faroese Case Study

Cross-lingual dependency parsing involves transferring syntactic knowledge from one language to another. It is a crucial component for inducing dependency parsers in low-resource scenarios where no training data for a language exists. Using Faroese as the target language, we compare two approaches using annotation projection: first, projecting from multiple monolingual source models; second, projecting from a single polyglot model which is trained on the combination of all source languages. Furthermore, we reproduce multi-source projection (Tyers et al., 2018), in which dependency trees of multiple sources are combined. Finally, we apply multi-treebank modelling to the projected treebanks, in addition to or alternatively to polyglot modelling on the source side. We find that polyglot training on the source languages produces an overall trend of better results on the target language but the single best result for the target language is obtained by projecting from monolingual source parsing models and then training multi-treebank POS tagging and parsing models on the target side.

cs.CL

Hindered nematic alignment of hematite spindles in viscoelastic matrices

The viscoelastic behavior of composites consisting of spindle-shaped hematite particles in poly-N-isopropylacrylamide hydrogels is investigated both, by means of rheological oscillatory shear experiments, and the field-induced alignment of these mesoscale, anisotropic particles in external magnetic fields. Due to their magnetic moment and magnetic anisotropy hematite spindles align with their long axis perpendicular to the direction of an external magnetic field. The field induced torque acting on the magnetic particles leads to an elastic deformation of the hydrogel matrix. Thus, the field-dependent orientational distribution functions of anisotropic particles acting as microrheological probes depend on the elastic modulus of the hydrogel matrix. The orientational distribution functions are determined by means of Small Angle X-ray Scattering experiments in presence of external magnetic fields. With increasing elasticity of the hydrogels, tuned via the polymer volume fraction and the crosslinking density, the field-induced alignment of these anisotropic, magnetic particles is progressively hindered. The microrheological results are in accordance to macrorheological experiments indicating increasing elasticity with increasing flux density of an external field.

cond-mat.soft

Depolarized light scattering from prolate anisotropic particles: the influence of the particle shape on the field autocorrelation function

We provide a theoretical analysis for the intermediate scattering function typically measured in depolarized dynamic light scattering experiments. We calculate the field autocorrelation function $g_1^{\rm VH}(Q,t)$ in dependence on the wave vector $Q$ and the time $t$ explicitly in a vertical-horizontal scattering geometry for differently shaped solids of revolution. The shape of prolate cylinders, spherocylinders, spindles, and double cones with variable aspect ratio is expanded in rotational invariants $f_{lm}(r)$. By Fourier transform of these expansion coefficients, a formal multipole expansion of the scattering function is obtained, which is used to calculate the weighting coefficients appearing in the depolarized scattering function. In addition to translational and rotational diffusion, especially the translational-rotational coupling of shape-anisotropic objects is considered. From the short-time behavior of the intermediate scattering function, the first cumulants $\Gamma(Q)$ are calculated. In a depolarized scattering experiment, they deviate from the simple proportionality to $Q^2$. The coefficients $f_{lm}(Q)$ strongly depend on the geometry and aspect ratio of the particles. The time dependence, in addition, is governed by the translational and rotational diffusion tensors, which are calculated by means of bead models for differently shaped particles in dependence on their aspect ratio. Therefore, our analysis shows how details of the particle shape---beyond their aspect ratio---can be determined by a precise scattering experiment. This is of high relevance in understanding smart materials which involve suspensions of anisotropic colloidal particles.

cond-mat.soft