SearcharxivSearch

arXiv subjects

Nicki Holighaus

Publications and source records attributed to Nicki Holighaus.

At least 19 recordsLinked to original sources

Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown a strong denoising performance in speech and music processing. However, their direct application to bioacoustic signals is limited by the scarcity of clean training data. To address this issue, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training sets and to design a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalizations (USVs) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In actual field recordings, the proposed method enhances fundamental and harmonic partial ridge tracking compared to our previous signal-processing approach. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. Our proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. Although we focus on USVs, the proposed approach should be broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridgebased training set synthesis and denoising.

cs.SD

SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform

Phase vocoder-based time-stretching is a widely used technique for the time-scale modification of audio signals. However, conventional implementations suffer from ``percussion smearing,'' a well-known artifact that significantly degrades the quality of percussive components. We attribute this artifact to a fundamental time-scale mismatch between the temporally smeared magnitude spectrogram and the localized, newly generated phase. To address this, we propose SELEBI, a signal-adaptive phase vocoder algorithm that significantly reduces percussion smearing while preserving stability and the perfect reconstruction property. Unlike conventional methods that rely on heuristic processing or component separation, our approach leverages the nonstationary Gabor transform. By dynamically adapting analysis window lengths to assign short windows to intervals containing significant energy associated with percussive components, we directly compute a temporally localized magnitude spectrogram from the time-domain signal. This approach ensures greater consistency between the temporal structures of the magnitude and phase. Furthermore, the perfect reconstruction property of the nonstationary Gabor transform guarantees stable, high-fidelity signal synthesis, in contrast to previous heuristic approaches. Experimental results demonstrate that the proposed method effectively mitigates percussion smearing and yields natural sound quality.

eess.AS

Aliasing in Convnets: A Frame-Theoretic Perspective

Using a stride in a convolutional layer inherently introduces aliasing, which has implications for numerical stability and statistical generalization. While techniques such as the parametrizations via paraunitary systems have been used to promote orthogonal convolution and thus ensure Parseval stability, a general analysis of aliasing and its effects on the stability has not been done in this context. In this article, we adapt a frame-theoretic approach to describe aliasing in convolutional layers with 1D kernels, leading to practical estimates for stability bounds and characterizations of Parseval stability, that are tailored to take short kernel sizes into account. From this, we derive two computationally very efficient optimization objectives that promote Parseval stability via systematically suppressing aliasing. Finally, for layers with random kernels, we derive closed-form expressions for the expected value and variance of the terms that describe the aliasing effects, revealing fundamental insights into the aliasing behavior at initialization.

cs.LG

Discretization of Continuous Frames by Quasi-Monte Carlo Methods

We introduce a discretization scheme for continuous localized frames using quasi-Monte Carlo integration and discrepancy theory. By generalizing classical concepts, we define a discrepancy measure on the entire phase space $\mathbb{R}^2$ and establish a corresponding Koksma-Hlawka inequality. This approach enables control over the density of the discretized frame and ensures the universality of the sampling set, relying only on the discrepancy of the sampling set and on the Sobolev-type seminorm of an iterated kernel rather than on specific frame properties.

math.FA

ISAC: An Invertible and Stable Auditory Filter Bank with Customizable Kernels for ML Integration

This paper introduces ISAC, an invertible and stable, perceptually-motivated filter bank that is specifically designed to be integrated into machine learning paradigms. More precisely, the center frequencies and bandwidths of the filters are chosen to follow a non-linear, auditory frequency scale, the filter kernels have user-defined maximum temporal support and may serve as learnable convolutional kernels, and there exists a corresponding filter bank such that both form a perfect reconstruction pair. ISAC provides a powerful and user-friendly audio front-end suitable for any application, including analysis-synthesis schemes.

cs.SD

Smoothness spaces for warped time-frequency representations -- Decomposition spaces and embedding relations

In a recent paper, we have shown that warped time-frequency representations provide a rich framework for the construction and study of smoothness spaces matched to very general phase space geometries obtained by diffeomorphic deformations of $\mathbb{R}^d$. Here, we study these spaces, obtained through the application of general coorbit theory, using the framework of decomposition spaces. This allows us to derive embedding relations between coorbit spaces associated to different warping functions, and relate them to established, important smothness spaces. In particular, we show that we obtain $α$-modulation spaces and spaces of dominating mixed smoothness as special cases and, in contrast, that this is only possible for Besov spaces if $d=1$.

math.FA

Coorbit theory of warped time-frequency systems in $\mathbb{R}^d$

Warped time-frequency systems have recently been introduced as a class of structured continuous frames for functions on the real line. Herein, we generalize this framework to the setting of functions of arbitrary dimensionality. After showing that the basic properties of warped time-frequency representations carry over to higher dimensions, we determine conditions on the warping function which guarantee that the associated Gramian is well-localized, so that associated families of coorbit spaces can be constructed. We then show that discrete Banach frame decompositions for these coorbit spaces can be obtained by sampling the continuous warped time-frequency systems. In particular, this implies that sparsity of a given function $f$ in the discrete warped time-frequency dictionary is equivalent to membership of $f$ in the coorbit space. We put special emphasis on the case of radial warping functions, for which the relevant assumptions simplify considerably.

math.FA

Rotated time-frequency lattices are sets of stable sampling for continuous wavelet systems

We provide an example for the generating matrix $A$ of a two-dimensional lattice $Γ= A\mathbb{Z}^2$, such that the following holds: For any sufficiently smooth and localized mother wavelet $ψ$, there is a constant $β(A,ψ)>0$, such that $βΓ\cap (\mathbb{R}\times\mathbb{R}^+)$ is a set of stable sampling for the wavelet system generated by $ψ$, for all $0<β\leq β(A,ψ)$. The result and choice of the generating matrix are loosely inspired by the studies of low discrepancy sequences and uniform distribution modulo $1$. In particular, we estimate the number of lattice points contained in any axis parallel rectangle of fixed area. This estimate is combined with a recent sampling result for continuous wavelet systems, obtained via the oscillation method of general coorbit theory.

math.FA

Continuous warped time-frequency representations - Coorbit spaces and discretization

We present a novel family of continuous, linear time-frequency transforms adaptable to a multitude of (nonlinear) frequency scales. Similar to classical time-frequency or time-scale representations, the representation coefficients are obtained as inner products with the elements of a continuously indexed family of time-frequency atoms. These atoms are obtained from a single prototype function, by means of modulation, translation and warping. By warping we refer to the process of nonlinear evaluation according to a bijective, increasing function, the warping function. Besides showing that the resulting integral transforms fulfill certain basic, but essential properties, such as continuity and invertibility, we will show that a large subclass of warping functions gives rise to families of generalized coorbit spaces, i.e. Banach spaces of functions whose representations possess a certain localization. Furthermore, we obtain sufficient conditions for subsampled warped time-frequency systems to form atomic decompositions and Banach frames. To this end, we extend results previously presented by Fornasier and Rauhut to a larger class of function systems via a simple, but crucial modification. The proposed method allows for great flexibility, but by choosing particular warping functions $Φ$ we also recover classical time-frequency representations, e.g. $Φ(t) = ct$ provides the short-time Fourier transform and $Φ(t) = log_a(t)$ provides wavelet transforms. This is illustrated by a number of examples provided in the manuscript.

math.FA

Time-frequency analysis on flat tori and Gabor frames in finite dimensions

We provide the foundations of a Hilbert space theory for the short-time Fourier transform (STFT) where the flat tori \begin{equation*} \mathbb{T}_{N}^2=\mathbb{R}^2/(\mathbb{Z}\times N\mathbb{Z})=[0,1]\times \lbrack 0,N] \end{equation*} act as phase spaces. We work on an $N$-dimensional subspace $S_{N}$ of distributions periodic in time and frequency in the dual $S_0'(\mathbb{R})$ of the Feichtinger algebra $S_0(\mathbb{R})$ and equip it with an inner product. To construct the Hilbert space $S_{N}$ we apply a suitable double periodization operator to $S_0(\mathbb{R})$. On $S_{N}$, the STFT is applied as the usual STFT defined on $S_0'(\mathbb{R})$. This STFT is a continuous extension of the finite discrete Gabor transform from the lattice onto the entire flat torus. As such, sampling theorems on flat tori lead to Gabor frames in finite dimensions. For Gaussian windows, one is lead to spaces of analytic functions and the construction allows to prove a necessary and sufficient Nyquist rate type result, which is the analogue, for Gabor frames in finite dimensions, of a well known result of Lyubarskii and Seip-Wallst{é}n for Gabor frames with Gaussian windows and which, for $N$ odd, produces an explicit \emph{full spark Gabor frame}. The compactness of the phase space, the finite dimension of the signal spaces and our sampling theorem offer practical advantages in some applications. We illustrate this by discussing a problem of current research interest: recovering signals from the zeros of their noisy spectrograms.

math.FA

Grid-Based Decimation for Wavelet Transforms with Stably Invertible Implementation

The constant center frequency to bandwidth ratio (Q-factor) of wavelet transforms provides a very natural representation for audio data. However, invertible wavelet transforms have either required non-uniform decimation -- leading to irregular data structures that are cumbersome to work with -- or require excessively high oversampling with unacceptable computational overhead. Here, we present a novel decimation strategy for wavelet transforms that leads to stable representations with oversampling rates close to one and uniform decimation. Specifically, we show that finite implementations of the resulting representation are energy-preserving in the sense of frame theory. The obtained wavelet coefficients can be stored in a timefrequency matrix with a natural interpretation of columns as time frames and rows as frequency channels. This matrix structure immediately grants access to a large number of algorithms that are successfully used in time-frequency audio processing, but could not previously be used jointly with wavelet transforms. We demonstrate the application of our method in processing based on nonnegative matrix factorization, in onset detection, and in phaseless reconstruction.

eess.AS

Audio inpainting of music by means of neural networks

We studied the ability of deep neural networks (DNNs) to restore missing audio content based on its context, a process usually referred to as audio inpainting. We focused on gaps in the range of tens of milliseconds. The proposed DNN structure was trained on audio signals containing music and musical instruments, separately, with 64-ms long gaps. The input to the DNN was the context, i.e., the signal surrounding the gap, transformed into time-frequency (TF) coefficients. Our results were compared to those obtained from a reference method based on linear predictive coding (LPC). For music, our DNN significantly outperformed the reference method, demonstrating a generally good usability of the proposed DNN structure for inpainting complex audio signals like music.

cs.SD

Fast Matching Pursuit with Multi-Gabor Dictionaries

Finding the best K-sparse approximation of a signal in a redundant dictionary is an NP-hard problem. Suboptimal greedy matching pursuit (MP) algorithms are generally used for this task. In this work, we present an acceleration technique and an implementation of the matching pursuit algorithm acting on a multi-Gabor dictionary, i.e., a concatenation of several Gabor-type time-frequency dictionaries, each of which consisting of translations and modulations of a possibly different window and time and frequency shift parameters. The technique is based on pre-computing and thresholding inner products between atoms and on updating the residual directly in the coefficient domain, i.e., without the round-trip to the signal domain. Since the proposed acceleration technique involves an approximate update step, we provide theoretical and experimental results illustrating the convergence of the resulting algorithm. The implementation is written in C (compatible with C99 and C++11) and we also provide Matlab and GNU Octave interfaces. For some settings, the implementation is up to 70 times faster than the standard Matching Pursuit Toolkit (MPTK).

math.NA

Phase Vocoder Done Right

The phase vocoder (PV) is a widely spread technique for processing audio signals. It employs a short-time Fourier transform (STFT) analysis-modify-synthesis loop and is typically used for time-scaling of signals by means of using different time steps for STFT analysis and synthesis. The main challenge of PV used for that purpose is the correction of the STFT phase. In this paper, we introduce a novel method for phase correction based on phase gradient estimation and its integration. The method does not require explicit peak picking and tracking nor does it require detection of transients and their separate treatment. Yet, the method does not suffer from the typical phase vocoder artifacts even for extreme time stretching factors.

cs.SD

Audio Inpainting via $\ell_1$-Minimization and Dictionary Learning

Audio inpainting refers to signal processing techniques that aim at restoring missing or corrupted consecutive samples in audio signals. Prior works have shown that $\ell_1$- minimization with appropriate weighting is capable of solving audio inpainting problems, both for the analysis and the synthesis models. These models assume that audio signals are sparse with respect to some redundant dictionary and exploit that sparsity for inpainting purposes. Remaining within the sparsity framework, we utilize dictionary learning to further increase the sparsity and combine it with weighted $\ell_1$-minimization adapted for audio inpainting to compensate for the loss of energy within the gap after restoration. Our experiments demonstrate that our approach is superior in terms of signal-to-distortion ratio (SDR) and objective difference grade (ODG) compared with its original counterpart.

cs.SD

Non-iterative Filter Bank Phase (Re)Construction

Signal reconstruction from magnitude-only measurements presents a long-standing problem in signal processing. In this contribution, we propose a phase (re)construction method for filter banks with uniform decimation and controlled frequency variation. The suggested procedure extends the recently introduced phase-gradient heap integration and relies on a phase-magnitude relationship for filter bank coefficients obtained from Gaussian filters. Admissible filter banks are modeled as the discretization of certain generalized translation-invariant systems, for which we derive the phase-magnitude relationship explicitly. The implementation for discrete signals is described and the performance of the algorithm is evaluated on a range of real and synthetic signals.

cs.SD

Phase-Based Signal Representations for Scattering

The scattering transform is a non-linear signal representation method based on cascaded wavelet transform magnitudes. In this paper we introduce phase scattering, a novel approach where we use phase derivatives in a scattering procedure. We first revisit phase-related concepts for representing time-frequency information of audio signals, in particular, the partial derivatives of the phase in the time-frequency domain. By putting analytical and numerical results in a new light, we set the basis to extend the phase-based representations to higher orders by means of a scattering transform, which leads to well localized signal representations of large-scale structures. All the ideas are introduced in a general way and then applied using the STFT.

cs.SD

Time-Frequency Phase Retrieval for Audio -- The Effect of Transform Parameters

In audio processing applications, phase retrieval (PR) is often performed from the magnitude of short-time Fourier transform (STFT) coefficients. Although PR performance has been observed to depend on the considered STFT parameters and audio data, the extent of this dependence has not been systematically evaluated yet. To address this, we studied the performance of three PR algorithms for various types of audio content and various STFT parameters such as redundancy, time-frequency ratio, and the type of window. The quality of PR was studied in terms of objective difference grade and signal-to-noise ratio of the STFT magnitude, to provide auditory- and signal-based quality assessments. Our results show that PR quality improved with increasing redundancy, with a strong relevance of the time-frequency ratio. The effect of the audio content was smaller but still observable. The effect of the window was only significant for one of the PR algorithms. Interestingly, for a good PR quality, each of the three algorithms required a different set of parameters, demonstrating the relevance of individual parameter sets for a fair comparison across PR algorithms. Based on these results, we developed guidelines for optimizing STFT parameters for a given application.

eess.SP