SearcharxivSearch

arXiv subjects

Thomas Deppisch

Publications and source records attributed to Thomas Deppisch.

9 recordsLinked to original sources

Direction-Preserving MIMO Speech Enhancement Using a Neural Covariance Estimator

Multichannel speech enhancement is widely used as a front-end in microphone array processing systems. While most existing approaches produce a single enhanced signal, direction-preserving multiple-input multiple-output (MIMO) methods instead aim to provide enhanced multichannel signals that retain directional properties, enabling downstream applications such as beamforming, binaural rendering, and direction-of-arrival estimation. In this work, we propose a fully blind, direction-preserving MIMO speech enhancement method based on neural estimation of the spatial noise covariance matrix. A lightweight OnlineSpatialNet estimates a scale-normalized Cholesky factor of the frequency-domain noise covariance, which is combined with a direction-preserving MIMO Wiener filter to enhance speech while preserving the spatial characteristics of both target and residual noise. In contrast to prior approaches relying on oracle information or mask-based covariance estimation for single-output systems, the proposed method directly targets accurate multichannel covariance estimation with low computational complexity. Experimental results show improved speech enhancement, covariance estimation capability, and performance in downstream tasks over a mask-based baseline, approaching oracle performance with significantly fewer parameters and computational cost.

eess.AS

Residual Learning for Neural Ambisonics Encoders

Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provides a device-agnostic spatial audio representation by mapping array signals to spherical harmonic (SH) coefficients. In practice, however, accurate encoding remains challenging. While traditional linear encoders are signal-independent and robust, they amplify low-frequency noise and suffer from high-frequency spatial aliasing. On the other hand, neural network approaches can outperform linear encoders but they often assume idealized microphones and may perform inconsistently in real-world scenarios. To leverage their complementary strengths, we introduce a residual-learning framework that refines a linear encoder with corrections from a neural network. Using measured array transfer functions from smartglasses, we compare a UNet-based encoder from the literature with a new recurrent attention model. Our analysis reveals that both neural encoders only consistently outperform the linear baseline when integrated within the residual learning framework. In the residual configuration, both neural models achieve consistent and significant improvements across all tested metrics for in-domain data and moderate gains for out-of-domain data. Yet, coherence analysis indicates that all neural encoder configurations continue to struggle with directionally accurate high-frequency encoding.

eess.AS

Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices.

cs.SD

Blind Identification of Binaural Room Impulse Responses from Smart Glasses

Smart glasses are increasingly recognized as a key medium for augmented reality, offering a hands-free platform with integrated microphones and non-ear-occluding loudspeakers to seamlessly mix virtual sound sources into the real-world acoustic scene. To convincingly integrate virtual sound sources, the room acoustic rendering of the virtual sources must match the real-world acoustics. Information about a user's acoustic environment however is typically not available. This work uses a microphone array in a pair of smart glasses to blindly identify binaural room impulse responses (BRIRs) from a few seconds of speech in the real-world environment. The proposed method uses dereverberation and beamforming to generate a pseudo reference signal that is used by a multichannel Wiener filter to estimate room impulse responses which are then converted to BRIRs. The multichannel room impulse responses can be used to estimate room acoustic parameters which is shown to outperform baseline algorithms in the estimation of reverberation time and direct-to-reverberant energy ratio. Results from a listening experiment further indicate that the estimated BRIRs often reproduce the real-world room acoustics perceptually more convincingly than measured BRIRs from other rooms of similar size.

eess.AS

Direct and Residual Subspace Decomposition of Spatial Room Impulse Responses

Psychoacoustic experiments have shown that directional properties of the direct sound, salient reflections, and the late reverberation of an acoustic room response can have a distinct influence on the auditory perception of a given room. Spatial room impulse responses (SRIRs) capture those properties and thus are used for direction-dependent room acoustic analysis and virtual acoustic rendering. This work proposes a subspace method that decomposes SRIRs into a direct part, which comprises the direct sound and the salient reflections, and a residual, to facilitate enhanced analysis and rendering methods by providing individual access to these components. The proposed method is based on the generalized singular value decomposition and interprets the residual as noise that is to be separated from the other components of the reverberation. Large generalized singular values are attributed to the direct part, which is then obtained as a low-rank approximation of the SRIR. By advancing from the end of the SRIR toward the beginning while iteratively updating the residual estimate, the method adapts to spatio-temporal variations of the residual. The method is evaluated using a spatio-spectral error measure and simulated SRIRs of different rooms, microphone arrays, and ratios of direct sound to residual energy. The proposed method creates lower errors than existing approaches in all tested scenarios, including a scenario with two simultaneous reflections. A case study with measured SRIRs shows the applicability of the method under real-world acoustic conditions. A reference implementation is provided.

eess.AS

$\texttt{RGE++}:$ A $\texttt{C++}$ library to solve renormalisation group equations in quantum field theory

In recent years three-, four- and five-loop beta functions have been computed for various phenomenologically interesting models. However, most of these results have not been implemented in easy to use software packages. $\texttt{RGE++}$ bridges this gap by providing a flexible, template-based, $\texttt{C++}$ library to solve renormalisation group equations. Furthermore, we implement the available beta functions for the Standard Model, the minimal supersymmetric extension of the Standard Model and two-Higgs-doublet models, as well as right-handed neutrino extensions of the former two.

hep-ph

Little hierarchies solve the little fine-tuning problem: a case study in supersymmetry with heavy guinos

Radiative corrections with new heavy particles coupling to Higgs doublets destabilize the electroweak scale and require an ad-hoc counterterm cancelling the large loop contribution. If the mass scale m1 of these new particles in in the TeV range, this feature constitutes the "little fine-tuning problem". We consider the case that the new-physics spectrum has a little hierarchy with two particle mass scales m1, m2 and m2 = O(10 m1) and no tree-level couplings of the heavier particles to Higgs doublets. As a concrete example we study the (next-to-)minimal supersymmetric standard model ((N)MSSM) for the case that the gluino mass M3 is significantly larger than the stop mass parameters m_{L,R} and show that the usual one-loop fine-tuning analysis breaks down. If m_{L,R} is defined in the dimensional-reduction (DR-bar) or any other fundamental scheme, corrections enhanced by powers of M3^2/m_{L,R}^2 occur in all higher loop orders. After resumming these terms we find the fine-tuning measure substantially improved compared to the usual analyses with M3 <~ m_{L,R}. In our hierarchical scenario the stop self-energies grow like M3^2, so that the stop masses m_{L,R}^{OS} in the on-shell (OS) scheme are naturally much larger than their DR-bar counterparts m_{L,R}^{DR-bar}. This feature permits a novel solution to the little fine-tuning problem: DR-bar stop masses are close to the electroweak scale, but radiative corrections involving the heavy gluino push the OS masses, which are probed in collider searches, above their experimental lower limits. As a byproduct, we clarify which renormalization scheme must be used for squark masses in loop corrections to low-energy quantities such as the B-B-bar mixing amplitude.

hep-ph

Confronting SUSY SO(10) with updated Lattice and Neutrino Data

We present an updated fit of supersymmetric SO(10) models to quark and lepton masses and mixing parameters. Including latest results from lattice QCD determinations of quark masses and neutrino oscillation data, we show that fits neglecting supersymmetric threshold corrections are strongly disfavoured in our setup. Only when we include these corrections we find good fit points. We present $χ^2$-profiles for the threshold parameters, which show that in our setup the thresholds related to the third generation of fermions exhibit two rather narrow minima.

hep-ph

E6Tensors: A Mathematica Package for E6 Tensors

We present the Mathematica package E6Tensors, a tool for explicit tensor calculations in E6 gauge theories. In addition to matrix expressions for the group generators of E6, it provides structure constants, various higher rank tensors and expressions for the representations 27, 78, 351 and 351'. This paper comes along with a short manual including physically relevant examples. I further give a complete list of gauge invariant, renormalisable terms for superpotentials and Lagrangians.

hep-ph