SearcharxivSearch

arXiv subjects

Peter Balazs

Publications and source records attributed to Peter Balazs.

At least 19 recordsLinked to original sources

Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown a strong denoising performance in speech and music processing. However, their direct application to bioacoustic signals is limited by the scarcity of clean training data. To address this issue, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training sets and to design a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalizations (USVs) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In actual field recordings, the proposed method enhances fundamental and harmonic partial ridge tracking compared to our previous signal-processing approach. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. Our proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. Although we focus on USVs, the proposed approach should be broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridgebased training set synthesis and denoising.

cs.SD

Aliasing in Convnets: A Frame-Theoretic Perspective

Using a stride in a convolutional layer inherently introduces aliasing, which has implications for numerical stability and statistical generalization. While techniques such as the parametrizations via paraunitary systems have been used to promote orthogonal convolution and thus ensure Parseval stability, a general analysis of aliasing and its effects on the stability has not been done in this context. In this article, we adapt a frame-theoretic approach to describe aliasing in convolutional layers with 1D kernels, leading to practical estimates for stability bounds and characterizations of Parseval stability, that are tailored to take short kernel sizes into account. From this, we derive two computationally very efficient optimization objectives that promote Parseval stability via systematically suppressing aliasing. Finally, for layers with random kernels, we derive closed-form expressions for the expected value and variance of the terms that describe the aliasing effects, revealing fundamental insights into the aliasing behavior at initialization.

cs.LG

The Lifting Property for Frame Multipliers and Toeplitz Operators

Frame multipliers are an abstract version of Toeplitz operators in frame theory and consist of a composition of a multiplication operator with the analysis and synthesis operators. Whereas the boundedness properties of frame multipliers on Banach spaces associated to a frame, so-called coorbit spaces, are well understood, their invertibility is much more difficult. We show that frame multipliers with a positive symbol are Banach space isomorphisms between the corresponding coorbit spaces. The results resemble the lifting theorems in the theory of Besov spaces and modulation spaces. Indeed, the application of the abstract lifting theorem to Gabor frames yields a new lifting theorem between modulation spaces. A second application to Fock spaces yields isomorphisms between weighted Fock spaces. The main techniques are the theory of localized frames and existence of inverse-closed matrix algebras.

math.FA

Localized frames without inequalities

We consider countable families of vectors in a separable Hilbert space, which are mutually localized with respect to a fixed localized Riesz basis. We prove the equivalence of the frame property and nine conditions that do not involve any inequalities. This is done by studying the properties of their frame-related operators on the co-orbit spaces generated by the reference Riesz basis. We apply our main result to the setting of shift-invariant spaces and obtain new conditions for stable sets of sampling.

math.FA

ISAC: An Invertible and Stable Auditory Filter Bank with Customizable Kernels for ML Integration

This paper introduces ISAC, an invertible and stable, perceptually-motivated filter bank that is specifically designed to be integrated into machine learning paradigms. More precisely, the center frequencies and bandwidths of the filters are chosen to follow a non-linear, auditory frequency scale, the filter kernels have user-defined maximum temporal support and may serve as learnable convolutional kernels, and there exists a corresponding filter bank such that both form a perfect reconstruction pair. ISAC provides a powerful and user-friendly audio front-end suitable for any application, including analysis-synthesis schemes.

cs.SD

Localization of operator-valued frames

We introduce a localization concept for operator-valued frames, where the quality of localization is measured by the associated operator-valued Gram matrix belonging to some suitable Banach algebra. We prove that intrinsic localization of an operator-valued frame is preserved by its canonical dual. Moreover, we show that the series associated to the perfect reconstruction of an operator-valued frame converges not only in the underlying Hilbert space, but also in a whole class of associated (quasi-)Banach spaces. Finally, we apply our results to irregular Gabor g-frames.

math.FA

Localised frames for tensor product spaces

In this paper, we investigate whether the tensor product of two frames, each individually localised with respect to a spectral matrix algebra, is also localised with respect to a suitably chosen tensor product algebra. We provide a partial answer by constructing an involutive Banach algebra of rank-four tensors that is built from two solid spectral matrix algebras. We show that this algebra is inverse-closed, given that the original algebras satisfy a specific property related to operator-valued versions of these algebras. This condition is satisfied by all commonly used solid spectral matrix algebras. We then prove that the tensor product of two self-localised frames remains self-localised with respect to our newly constructed tensor algebra. Additionally, we discuss generalisations to localised frames of Hilbert-Schmidt operators, which may not necessarily consist of rank-one operators.

math.FA

Details on the distribution co-orbit space $\mathcal{H}^{\infty}_w$

Associated with every separable Hilbert space $\mathcal{H}$ and a given localized frame, there exists a natural test function Banach space $\mathcal{H}^1$ and a Banach distribution space $\mathcal{H}^{\infty}$ so that $\mathcal{H}^1 \subset \mathcal{H} \subset \mathcal{H}^{\infty}$. In this article we close some gaps in the literature and rigorously introduce the space $\mathcal{H}^{\infty}$ and its weighted variants $\mathcal{H}_w^{\infty}$ in a slightly more general setting and discuss some of their properties. In particular, we compare the underlying weak$^*$- with the norm topology associated with $\mathcal{H}_w^{\infty}$ and show that $(\mathcal{H}_w^{\infty}, \Vert \cdot \Vert_{\mathcal{H}_w^{\infty}})$ is a Banach space.

math.FA

On the inverse-closedness of operator-valued matrices with polynomial off-diagonal decay

We give a self-contained proof of a recently established $\mathcal{B}(\mathcal{H})$-valued version of Jaffards Lemma. That is, we show that the Jaffard algebra of $\mathcal{B}(\mathcal{H})$-valued matrices, whose operator norms of their respective entries decay polynomially off the diagonal, is a Banach algebra which is inverse-closed in the Banach algebra $\mathcal{B}(\ell^2(X;\mathcal{H}))$ of all bounded linear operators on $\ell^2(X;\mathcal{H})$, the Bochner-space of square-summable $\mathcal{H}$-valued sequences.

math.FA

Construction of generalized samplets in Banach spaces

Recently, samplets have been introduced as localized discrete signed measures which are tailored to an underlying data set. Samplets exhibit vanishing moments, i.e., their measure integrals vanish for all polynomials up to a certain degree, which allows for feature detection and data compression. In the present article, we extend the different construction steps of samplets to functionals in Banach spaces more general than point evaluations. To obtain stable representations, we assume that these functionals form frames with square-summable coefficients or even Riesz bases with square-summable coefficients. In either case, the corresponding analysis operator is injective and we obtain samplet bases with the desired properties by means of constructing an isometry of the analysis operator's image. Making the assumption that the dual of the Banach space under consideration is imbedded into the space of compactly supported distributions, the multilevel hierarchy for the generalized samplet construction is obtained by spectral clustering of a similarity graph for the functionals' supports. Based on this multilevel hierarchy, generalized samplets exhibit vanishing moments with respect to a given set of primitives within the Banach space. We derive an abstract localization result for the generalized samplet coefficients with respect to the samplets' support sizes and the approximability of the Banach space elements by the chosen primitives. Finally, we present three examples showcasing the generalized samplet framework.

math.FA

(Almost) Smooth Sailing: Towards Numerical Stability of Neural Networks Through Differentiable Regularization of the Condition Number

Maintaining numerical stability in machine learning models is crucial for their reliability and performance. One approach to maintain stability of a network layer is to integrate the condition number of the weight matrix as a regularizing term into the optimization algorithm. However, due to its discontinuous nature and lack of differentiability the condition number is not suitable for a gradient descent approach. This paper introduces a novel regularizer that is provably differentiable almost everywhere and promotes matrices with low condition numbers. In particular, we derive a formula for the gradient of this regularizer which can be easily implemented and integrated into existing optimization algorithms. We show the advantages of this approach for noisy classification and denoising of MNIST images.

cs.LG

Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement

Convolutional layers with 1-D filters are often used as frontend to encode audio signals. Unlike fixed time-frequency representations, they can adapt to the local characteristics of input data. However, 1-D filters on raw audio are hard to train and often suffer from instabilities. In this paper, we address these problems with hybrid solutions, i.e., combining theory-driven and data-driven approaches. First, we preprocess the audio signals via a auditory filterbank, guaranteeing good frequency localization for the learned encoder. Second, we use results from frame theory to define an unsupervised learning objective that encourages energy conservation and perfect reconstruction. Third, we adapt mixed compressed spectral norms as learning objectives to the encoder coefficients. Using these solutions in a low-complexity encoder-mask-decoder model significantly improves the perceptual evaluation of speech quality (PESQ) in speech enhancement.

cs.SD

Wiener pairs of Banach algebras of operator-valued matrices

In this article we introduce several new examples of Wiener pairs $\mathcal{A} \subseteq \mathcal{B}$, where $\mathcal{B} = \mathcal{B}(\ell^2(X;\mathcal{H}))$ is the Banach algebra of bounded operators acting on the Hilbert space-valued Bochner sequence space $\ell^2(X;\mathcal{H})$ and $\mathcal{A} = \mathcal{A}(X)$ is a Banach algebra consisting of operator-valued matrices indexed by some relatively separated set $X \subset \mathbb{R}^d$. In particular, we introduce $\mathcal{B}(\mathcal{H})$-valued versions of the Jaffard algebra, of certain weighted Schur-type algebras, of Banach algebras which are defined by more general off-diagonal decay conditions than polynomial decay, of weighted versions of the Baskakov-Gohberg-Sj\"ostrand algebra, and of anisotropic variations of all of these matrix algebras, and show that they are inverse-closed in $\mathcal{B}(\ell^2(X;\mathcal{H}))$. In addition, we obtain that each of these Banach algebras is symmetric.

math.FA

Injectivity of ReLU-layers: Tools from Frame Theory

Injectivity is the defining property of a mapping that ensures no information is lost and any input can be perfectly reconstructed from its output. By performing hard thresholding, the ReLU function naturally interferes with this property, making the injectivity analysis of ReLU layers in neural networks a challenging yet intriguing task that has not yet been fully solved. This article establishes a frame theoretic perspective to approach this problem. The main objective is to develop a comprehensive characterization of the injectivity behavior of ReLU layers in terms of all three involved ingredients: (i) the weights, (ii) the bias, and (iii) the domain where the data is drawn from. Maintaining a focus on practical applications, we limit our attention to bounded domains and present two methods for numerically approximating a maximal bias for given weights and data domains. These methods provide sufficient conditions for the injectivity of a ReLU layer on those domains and yield a novel practical methodology for studying the information loss in ReLU layers. Finally, we derive explicit reconstruction formulas based on the duality concept from frame theory.

cs.LG

Kernel theorems for operators on co-orbit spaces associated with localised frames

Kernel theorems, in general, provide a convenient representation of bounded linear operators. For the operator acting on a concrete function space, this means that its action on any element of the space can be expressed as a generalised integral operator, in a way reminiscent of the matrix representation of linear operators acting on finite dimensional vector spaces. We prove kernel theorems for bounded linear operators acting on co-orbit spaces associated with localised frames. Our two main results consist in characterising the spaces of operators whose generalised integral kernels belong to the co-orbit spaces of test functions and distributions associated with the tensor product of the localised frames respectively. Moreover, using a version of Schur's test, we establish a characterisation of the bounded linear operators between some specific co-orbit spaces.

math.FA

Weighted frames, weighted lower semi frames and unconditionally convergent multipliers

In this paper we ask when it is possible to transform a given sequence into a frame or a lower semi frame by multiplying the elements by numbers. In other words, we ask when a given sequence is a weighted frame or a weighted lower semi frame and for each case we formulate a conjecture. We determine several conditions under which these conjectures are true. Finally, we prove an equivalence between two older conjectures, the first one being that any unconditionally convergent multiplier can be written as a multiplier of Bessel sequences by shifting of weights, and the second one that every unconditionally convergent multiplier which is invertible can be written as a multiplier of frames by shifting of weights. We also show that these conjectures are also related to one of the newly posed conjectures.

math.FA

An unbounded operator theory approach to lower frame and Riesz-Fischer sequences

Frames and orthonormal bases are naturally linked to bounded operators. To tackle unbounded operators those sequences might not be well suited. This has already been noted by von Neumann in the 1920ies. But modern frame theory also investigates other sequences, including those that are not naturally linked to bounded operators. The focus of this manuscript will be two such kind of sequences: lower frame and Riesz-Fischer sequences. We will discuss the inter-relation of those sequences. We will fill a hole existing in the literature regarding the classification of those sequences by their synthesis operator. We will use the idea of generalized frame operator and Gram matrix and extend it. We will use that to show properties for canonical duals for lower frame sequences, like e.g. a minimality condition regarding its coefficients. We will also show that other results that are known for frames can be generalized to lower frame sequences. To be able to tackle these tasks, we had to revisit the concept of invertibility (in particular for non-closed operators). In addition, we are able to define a particular adjoint, which is uniquely defined for any operator.

math.FA

Instabilities in Convnets for Raw Audio

What makes waveform-based deep learning so hard? Despite numerous attempts at training convolutional neural networks (convnets) for filterbank design, they often fail to outperform hand-crafted baselines. These baselines are linear time-invariant systems: as such, they can be approximated by convnets with wide receptive fields. Yet, in practice, gradient-based optimization leads to suboptimal approximations. In our article, we approach this phenomenon from the perspective of initialization. We present a theory of large deviations for the energy response of FIR filterbanks with random Gaussian weights. We find that deviations worsen for large filters and locally periodic input signals, which are both typical for audio signal processing applications. Numerical simulations align with our theory and suggest that the condition number of a convolutional layer follows a logarithmic scaling law between the number and length of the filters, which is reminiscent of discrete wavelet bases.

cs.LG