SearcharxivSearch

arXiv subjects

Bertram Taetz

Publications and source records attributed to Bertram Taetz.

14 recordsLinked to original sources

Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation

Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M). While foundational models have achieved strong performance in static settings, autonomous agents operating in dynamic environments must continuously incorporate new motion concepts -- such as novel athletic styles or specialized gestures -- without catastrophic forgetting of previously acquired skills. We investigate the stability-plasticity trade-off in bidirectional motion-language learning under sequential task exposure. Building on a frozen large language model backbone, we introduce low-rank adaptation (LoRA) variants designed to mitigate inter-task interference. We specifically propose mixture-of-experts architectures that utilize an autoencoder-based router to select task-specific experts at inference time, so that no task-label is needed. To evaluate these methods, we establish a reproducible five-task benchmark derived from HumanML3D through semantic clustering of motion descriptions. Our experimental results demonstrate near-zero forgetting across both M2T and T2M directions while maintaining high generation and captioning quality. Furthermore, we show that hard expert selection via routing significantly outperforms soft expert blending in quality metrics, indicating that preserving expert isolation is critical for maintaining performance in our continual learning setting. Finally, we observe that a divergence between token-level accuracy and downstream generation quality may occur, highlighting the need for more comprehensive evaluation protocols in future research on lifelong motion-language agents.

cs.LG

Amortized Inverse Kinematics via Graph Attention for Real-Time Human Avatar Animation

Inverse kinematics (IK) is a core operation in animation, robotics, and biomechanics: given Cartesian constraints, recover joint rotations under a known kinematic tree. In many real-time human avatar pipelines, the available signal per frame is a sparse set of tracked 3D joint positions, whereas animation systems require joint orientations to drive skinning. Recovering full orientations from positions is underconstrained, most notably because twist about bone axes is ambiguous, and classical IK solvers typically rely on iterative optimization that can be slow and sensitive to noisy inputs. We introduce IK-GAT, a lightweight graph-attention network that reconstructs full-body joint orientations from 3D joint positions in a single forward pass. The model performs message passing over the skeletal parent-child graph to exploit kinematic structure during rotation inference. To simplify learning, IK-GAT predicts rotations in a bone-aligned world-frame representation anchored to rest-pose bone frames. This parameterization makes the twist axis explicit and is exactly invertible to standard parent-relative local rotations given the kinematic tree and rest pose. The network uses a continuous 6D rotation representation and is trained with a geodesic loss on SO(3) together with an optional forward-kinematics consistency regularizer. IK-GAT produces animation-ready local rotations that can directly drive a rigged avatar or be converted to pose parameters of SMPL-like body models for real-time and online applications. With 374K parameters and over 650 FPS on CPU, IK-GAT outperforms VPoser-based per-frame iterative optimization without warm-start at significantly lower cost, and is robust to initial pose and input noise

cs.CV

Continual Learning for Image Captioning through Improved Image-Text Alignment

Generating accurate and coherent image captions in a continual learning setting remains a major challenge due to catastrophic forgetting and the difficulty of aligning evolving visual concepts with language over time. In this work, we propose a novel multi-loss framework for continual image captioning that integrates semantic guidance through prompt-based continual learning and contrastive alignment. Built upon a pretrained ViT-GPT-2 backbone, our approach combines standard cross-entropy loss with three additional components: (1) a prompt-based cosine similarity loss that aligns image embeddings with synthetically constructed prompts encoding objects, attributes, and actions; (2) a CLIP-style loss that promotes alignment between image embeddings and target caption embedding; and (3) a language-guided contrastive loss that employs a triplet loss to enhance class-level discriminability between tasks. Notably, our approach introduces no additional overhead at inference time and requires no prompts during caption generation. We find that this approach mitigates catastrophic forgetting, while achieving better semantic caption alignment compared to state-of-the-art methods. The code can be found via the following link: https://github.com/Gepardius/Taetz_Bordelius_Continual_ImageCaptioning.

cs.CV

MinJointTracker: Real-time inertial kinematic chain tracking with joint position estimation and minimal state size

Inertial motion capture is a promising approach for capturing motion outside the laboratory. However, as one major drawback, most of the current methods require different quantities to be calibrated or computed offline as part of the setup process, such as segment lengths, relative orientations between inertial measurement units (IMUs) and segment coordinate frames (IMU-to-segment calibrations) or the joint positions in the IMU frames. This renders the setup process inconvenient. This work contributes to real-time capable calibration-free inertial tracking of a kinematic chain, i.e. simultaneous recursive Bayesian estimation of global IMU angular kinematics and joint positions in the IMU frames, with a minimal state size. Experimental results on simulated IMU data from a three-link kinematic chain (manipulator study) as well as re-simulated IMU data from healthy humans walking (lower body study) show that the calibration-free and lightweight algorithm provides not only drift-free relative but also drift-free absolute orientation estimates with a global heading reference for only one IMU as well as robust and fast convergence of joint position estimates in the different movement scenarios.

cs.RO

Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary

This article explores a critical gap in Mixed Reality (MR) technology: while advances have been made, MR still struggles to authentically replicate human embodiment and socio-motor interaction. For MR to enable truly meaningful social experiences, it needs to incorporate multi-modal data streams and multi-agent interaction capabilities. To address this challenge, we present a comprehensive glossary covering key topics such as Virtual Characters and Autonomisation, Responsible AI, Ethics by Design, and the Scientific Challenges of Social MR within Neuroscience, Embodiment, and Technology. Our aim is to drive the transformative evolution of MR technologies that prioritize human-centric innovation, fostering richer digital connections. We advocate for MR systems that enhance social interaction and collaboration between humans and virtual autonomous agents, ensuring inclusivity, ethical design and psychological safety in the process.

cs.HC

Autoencoder Attractors for Uncertainty Estimation

The reliability assessment of a machine learning model's prediction is an important quantity for the deployment in safety critical applications. Not only can it be used to detect novel sceneries, either as out-of-distribution or anomaly sample, but it also helps to determine deficiencies in the training data distribution. A lot of promising research directions have either proposed traditional methods like Gaussian processes or extended deep learning based approaches, for example, by interpreting them from a Bayesian point of view. In this work we propose a novel approach for uncertainty estimation based on autoencoder models: The recursive application of a previously trained autoencoder model can be interpreted as a dynamical system storing training examples as attractors. While input images close to known samples will converge to the same or similar attractor, input samples containing unknown features are unstable and converge to different training samples by potentially removing or changing characteristic features. The use of dropout during training and inference leads to a family of similar dynamical systems, each one being robust on samples close to the training distribution but unstable on new features. Either the model reliably removes these features or the resulting instability can be exploited to detect problematic input samples. We evaluate our approach on several dataset combinations as well as on an industrial application for occupant classification in the vehicle interior for which we additionally release a new synthetic dataset.

cs.LG

Autoencoder for Synthetic to Real Generalization: From Simple to More Complex Scenes

Learning on synthetic data and transferring the resulting properties to their real counterparts is an important challenge for reducing costs and increasing safety in machine learning. In this work, we focus on autoencoder architectures and aim at learning latent space representations that are invariant to inductive biases caused by the domain shift between simulated and real images showing the same scenario. We train on synthetic images only, present approaches to increase generalizability and improve the preservation of the semantics to real datasets of increasing visual complexity. We show that pre-trained feature extractors (e.g. VGG) can be sufficient for generalization on images of lower complexity, but additional improvements are required for visually more complex scenes. To this end, we demonstrate a new sampling technique, which matches semantically important parts of the image, while randomizing the other parts, leads to salient feature extraction and a neglection of unimportant parts. This helps the generalization to real data and we further show that our approach outperforms fine-tuned classification models.

cs.CV

Autoencoder Based Inter-Vehicle Generalization for In-Cabin Occupant Classification

Common domain shift problem formulations consider the integration of multiple source domains, or the target domain during training. Regarding the generalization of machine learning models between different car interiors, we formulate the criterion of training in a single vehicle: without access to the target distribution of the vehicle the model would be deployed to, neither with access to multiple vehicles during training. We performed an investigation on the SVIRO dataset for occupant classification on the rear bench and propose an autoencoder based approach to improve the transferability. The autoencoder is on par with commonly used classification models when trained from scratch and sometimes out-performs models pre-trained on a large amount of data. Moreover, the autoencoder can transform images from unknown vehicles into the vehicle it was trained on. These results are corroborated by an evaluation on real infrared images from two vehicle interiors.

cs.CV

Illumination Normalization by Partially Impossible Encoder-Decoder Cost Function

Images recorded during the lifetime of computer vision based systems undergo a wide range of illumination and environmental conditions affecting the reliability of previously trained machine learning models. Image normalization is hence a valuable preprocessing component to enhance the models' robustness. To this end, we introduce a new strategy for the cost function formulation of encoder-decoder networks to average out all the unimportant information in the input images (e.g. environmental features and illumination changes) to focus on the reconstruction of the salient features (e.g. class instances). Our method exploits the availability of identical sceneries under different illumination and environmental conditions for which we formulate a partially impossible reconstruction target: the input image will not convey enough information to reconstruct the target in its entirety. Its applicability is assessed on three publicly available datasets. We combine the triplet loss as a regularizer in the latent space representation and a nearest neighbour search to improve the generalization to unseen illuminations and class instances. The importance of the aforementioned post-processing is highlighted on an automotive application. To this end, we release a synthetic dataset of sceneries from three different passenger compartments where each scenery is rendered under ten different illumination and environmental conditions: see https://sviro.kl.dfki.de

cs.CV

Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation

Modern large displacement optical flow algorithms usually use an initialization by either sparse descriptor matching techniques or dense approximate nearest neighbor fields. While the latter have the advantage of being dense, they have the major disadvantage of being very outlier-prone as they are not designed to find the optical flow, but the visually most similar correspondence. In this article we present a dense correspondence field approach that is much less outlier-prone and thus much better suited for optical flow estimation than approximate nearest neighbor fields. Our approach does not require explicit regularization, smoothing (like median filtering) or a new data term. Instead we solely rely on patch matching techniques and a novel multi-scale matching strategy. We also present enhancements for outlier filtering. We show that our approach is better suited for large displacement optical flow estimation than modern descriptor matching techniques. We do so by initializing EpicFlow with our approach instead of their originally used state-of-the-art descriptor matching technique. We significantly outperform the original EpicFlow on MPI-Sintel, KITTI 2012, KITTI 2015 and Middlebury. In this extended article of our former conference publication we further improve our approach in matching accuracy as well as runtime and present more experiments and insights.

cs.CV

Towards Self-Calibrating Inertial Body Motion Capture

This paper presents a novel online capable method for simultaneous estimation of human motion in terms of segment orientations and positions along with sensor-to-segment calibration parameters from inertial sensors attached to the body. In order to solve this ill-posed estimation problem, state-of-the-art motion, measurement and biomechanical models are combined with new stochastic equations and priors. These are based on the kinematics of multi-body systems, anatomical and body shape information, as well as, parameter properties for regularisation. This leads to a constrained weighted least squares problem that is solved in a sliding window fashion. Magnetometer information is currently only used for initialisation, while the estimation itself works without magnetometers. The method was tested on simulated, as well as, on real data, captured from a lower body configuration.

eess.SY

Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation

Modern large displacement optical flow algorithms usually use an initialization by either sparse descriptor matching techniques or dense approximate nearest neighbor fields. While the latter have the advantage of being dense, they have the major disadvantage of being very outlier prone as they are not designed to find the optical flow, but the visually most similar correspondence. In this paper we present a dense correspondence field approach that is much less outlier prone and thus much better suited for optical flow estimation than approximate nearest neighbor fields. Our approach is conceptually novel as it does not require explicit regularization, smoothing (like median filtering) or a new data term, but solely our novel purely data based search strategy that finds most inliers (even for small objects), while it effectively avoids finding outliers. Moreover, we present novel enhancements for outlier filtering. We show that our approach is better suited for large displacement optical flow estimation than state-of-the-art descriptor matching techniques. We do so by initializing EpicFlow (so far the best method on MPI-Sintel) with our Flow Fields instead of their originally used state-of-the-art descriptor matching technique. We significantly outperform the original EpicFlow on MPI-Sintel, KITTI and Middlebury.

cs.CV

A high-order unstaggered constrained transport method for the 3D ideal magnetohydrodynamic equations based on the method of lines

Numerical methods for solving the ideal magnetohydrodynamic (MHD) equations in more than one space dimension must confront the challenge of controlling errors in the discrete divergence of the magnetic field. One approach that has been shown successful in stabilizing MHD calculations are constrained transport (CT) schemes. CT schemes can be viewed as predictor-corrector methods for updating the magnetic field, where a magnetic field value is first predicted by a method that does not exactly preserve the divergence-free condition on the magnetic field, followed by a correction step that aims to control these divergence errors. In Helzel et al. (2011) the authors presented an unstaggered constrained transport method for the MHD equations on 3D Cartesian grids. In this work we generalize the method of Helzel et al. (2011) in three important ways: (1) we remove the need for operator splitting by switching to an appropriate method of lines discretization and coupling this with a non-conservative finite volume method for the magnetic vector potential equation, (2) we increase the spatial and temporal order of accuracy of the entire method to third order, and (3) we develop the method so that it is applicable on both Cartesian and logically rectangular mapped grids. The evolution equation for the magnetic vector potential is solved using a non-conservative finite volume method. The curl of the magnetic potential is computed via a third-order accurate discrete operator that is derived from appropriate application of the divergence theorem and subsequent numerical quadrature on element faces. Special artificial resistivity limiters are used to control unphysical oscillations in the magnetic potential and field components across shocks. Test computations are shown that confirm third order accuracy for smooth test problems and high-resolution for test problems with shock waves.

math.NA

An Unstaggered Constrained Transport Method for the 3D Ideal Magnetohydrodynamic Equations

Numerical methods for solving the ideal magnetohydrodynamic (MHD) equations in more than one space dimension must either confront the challenge of controlling errors in the discrete divergence of the magnetic field, or else be faced with nonlinear numerical instabilities. One approach for controlling the discrete divergence is through a so-called constrained transport method, which is based on first predicting a magnetic field through a standard finite volume solver, and then correcting this field through the appropriate use of a magnetic vector potential. In this work we develop a constrained transport method for the 3D ideal MHD equations that is based on a high-resolution wave propagation scheme. Our proposed scheme is the 3D extension of the 2D scheme developed by Rossmanith [SIAM J. Sci. Comp. 28, 1766 (2006)], and is based on the high-resolution wave propagation method of Langseth and LeVeque [J. Comp. Phys. 165, 126 (2000)]. In particular, in our extension we take great care to maintain the three most important properties of the 2D scheme: (1) all quantities, including all components of the magnetic field and magnetic potential, are treated as cell-centered; (2) we develop a high-resolution wave propagation scheme for evolving the magnetic potential; and (3) we develop a wave limiting approach that is applied during the vector potential evolution, which controls unphysical oscillations in the magnetic field. One of the key numerical difficulties that is novel to 3D is that the transport equation that must be solved for the magnetic vector potential is only weakly hyperbolic. In presenting our numerical algorithm we describe how to numerically handle this problem of weak hyperbolicity, as well as how to choose an appropriate gauge condition. The resulting scheme is applied to several numerical test cases.

math.NA