Searcharxiv⌕ Search

arXiv subjects

Günther Koliander

Publications and source records attributed to Günther Koliander.

At least 19 recordsLinked to original sources

Zeros of the Spectrogram of Colored Noise

We study the expected number of zeros of the Short-Time Fourier Transform (STFT) with a Gaussian window for signals degraded by complex Gaussian colored noise. We provide an exact formula for this quantity and investigate its asymptotic behavior. From a computational perspective, we propose an algorithmic method to recover a smoothed power spectral density (PSD) directly from the spatial distribution of the transform's zeros. Our results formalize the commonly-held heuristic that detection algorithms based on spectrogram zeros, though designed for white noise, also perform adequately under moderately colored noise.

math.PR↗

PhaseJumps: fast computation of zeros from planar grid samples

We consider complex-valued functions on the complex plane and the task of computing their zeros from samples taken along a finite grid. We introduce PhaseJumps, an algorithm based on comparing changes in the complex phase and local oscillations among neighboring grid points. The algorithm is applicable to possibly non-analytic input functions, and also computes the direction of phase winding around zeros. PhaseJumps provides a first effective means to compute the zeros of the short-time Fourier transform of an analog signal with respect to a general analysis window, and makes certain recent signal processing insights more widely applicable, overcoming previous constraints to analytic transformations. We study the performance of (a variant of) PhaseJumps under a stochastic input model motivated by signal processing applications and show that the input instances that may cause the algorithm to fail are fragile, in the sense that they are regularized by additive noise (smoothed analysis). Precisely, given samples of a function on a grid with spacing $δ$, we show that our algorithm computes zeros with accuracy $\sqrtδ$ in the Wasserstein metric with failure probability $O\big(\log^2(\tfrac{1}δ) δ\big)$, while numerical experiments suggest even better performance.

math.NA↗

Hyperuniformity and non-hyperuniformity of zeros of Gaussian Weyl-Heisenberg Functions

We study zero sets of twisted stationary Gaussian random functions on the complex plane, i.e., Gaussian random functions that are stochastically invariant under the action of the Weyl-Heisenberg group. This model includes translation-invariant Gaussian entire functions (GEFs), and also many other non-analytic examples, in which case winding numbers around zeros can be either positive or negative. We investigate zero statistics both when zeros are weighted with their winding numbers (charged zero set) and when they are not (uncharged zero set). We show that the variance of the charged zero statistic always grows linearly with the radius of the observation disk (hyperuniformity). Importantly, this holds for functions with possibly non-zero means and without assuming additional symmetries such as radiality. With respect to uncharged zero statistics, we provide an example for which the variance grows with the area of the observation disk (non-hyperuniformity). This is used to show that, while the zeros of GEFs are hyperuniform, the set of their critical points fails to be so. Our work contributes to recent developments in statistical signal processing, where the time-frequency profile of a non-stationary signal embedded into noise is revealed by performing a statistical test on the zeros of its spectrogram ("silent points"). We show that empirical spectrogram zero counts enjoy moderate deviations from their ensemble averages over large observation windows (something that was previously known only for pure noise). In contrast, we also show that spectrogram maxima ("loud points") fail to enjoy a similar property. This gives the first formal evidence for the statistical superiority of silent points over the competing feature of loud points, a fact that has been noted by practitioners.

math.PR↗

Lossless Analog Compression

We establish the fundamental limits of lossless analog compression by considering the recovery of arbitrary m-dimensional real random vectors x from the noiseless linear measurements y=Ax with n x m measurement matrix A. Our theory is inspired by the groundbreaking work of Wu and Verdu (2010) on almost lossless analog compression, but applies to the nonasymptotic, i.e., fixed-m case, and considers zero error probability. Specifically, our achievability result states that, for almost all A, the random vector x can be recovered with zero error probability provided that n > K(x), where K(x) is given by the infimum of the lower modified Minkowski dimension over all support sets U of x. We then particularize this achievability result to the class of s-rectifiable random vectors as introduced in Koliander et al. (2016); these are random vectors of absolutely continuous distribution -- with respect to the s-dimensional Hausdorff measure -- supported on countable unions of s-dimensional differentiable submanifolds of the m-dimensional real coordinate space. Countable unions of differentiable submanifolds include essentially all signal models used in the compressed sensing literature. Specifically, we prove that, for almost all A, s-rectifiable random vectors x can be recovered with zero error probability from n>s linear measurements. This threshold is, however, found not to be tight as exemplified by the construction of an s-rectifiable random vector that can be recovered with zero error probability from n<s linear measurements. This leads us to the introduction of the new class of s-analytic random vectors, which admit a strong converse in the sense of n greater than or equal to s being necessary for recovery with probability of error smaller than one. The central conceptual tools in the development of our theory are geometric measure theory and the theory of real analytic functions.

math.FA↗

Completion of Matrices with Low Description Complexity

We propose a theory for matrix completion that goes beyond the low-rank structure commonly considered in the literature and applies to general matrices of low description complexity. Specifically, complexity of the sets of matrices encompassed by the theory is measured in terms of Hausdorff and upper Minkowski dimensions. Our goal is the characterization of the number of linear measurements, with an emphasis on rank-$1$ measurements, needed for the existence of an algorithm that yields reconstruction, either perfect, with probability 1, or with arbitrarily small probability of error, depending on the setup. Concretely, we show that matrices taken from a set $\mathcal{U}$ such that $\mathcal{U}-\mathcal{U}$ has Hausdorff dimension $s$ can be recovered from $k>s$ measurements, and random matrices supported on a set $\mathcal{U}$ of Hausdorff dimension $s$ can be recovered with probability 1 from $k>s$ measurements. What is more, we establish the existence of recovery mappings that are robust against additive perturbations or noise in the measurements. Concretely, we show that there are $β$-Hölder continuous mappings recovering matrices taken from a set of upper Minkowski dimension $s$ from $k>2s/(1-β)$ measurements and, with arbitrarily small probability of error, random matrices supported on a set of upper Minkowski dimension $s$ from $k>s/(1-β)$ measurements. The numerous concrete examples we consider include low-rank matrices, sparse matrices, QR decompositions with sparse R-components, and matrices of fractal nature.

cs.IT↗

Rotated time-frequency lattices are sets of stable sampling for continuous wavelet systems

We provide an example for the generating matrix $A$ of a two-dimensional lattice $Γ= A\mathbb{Z}^2$, such that the following holds: For any sufficiently smooth and localized mother wavelet $ψ$, there is a constant $β(A,ψ)>0$, such that $βΓ\cap (\mathbb{R}\times\mathbb{R}^+)$ is a set of stable sampling for the wavelet system generated by $ψ$, for all $0<β\leq β(A,ψ)$. The result and choice of the generating matrix are loosely inspired by the studies of low discrepancy sequences and uniform distribution modulo $1$. In particular, we estimate the number of lattice points contained in any axis parallel rectangle of fixed area. This estimate is combined with a recent sampling result for continuous wavelet systems, obtained via the oscillation method of general coorbit theory.

math.FA↗

Lossy Compression of General Random Variables

This paper is concerned with the lossy compression of general random variables, specifically with rate-distortion theory and quantization of random variables taking values in general measurable spaces such as, e.g., manifolds and fractal sets. Manifold structures are prevalent in data science, e.g., in compressed sensing, machine learning, image processing, and handwritten digit recognition. Fractal sets find application in image compression and in the modeling of Ethernet traffic. Our main contributions are bounds on the rate-distortion function and the quantization error. These bounds are very general and essentially only require the existence of reference measures satisfying certain regularity conditions in terms of small ball probabilities. To illustrate the wide applicability of our results, we particularize them to random variables taking values in i) manifolds, namely, hyperspheres and Grassmannians, and ii) self-similar sets characterized by iterated function systems satisfying the weak separation property.

math.PR↗

Grid-Based Decimation for Wavelet Transforms with Stably Invertible Implementation

The constant center frequency to bandwidth ratio (Q-factor) of wavelet transforms provides a very natural representation for audio data. However, invertible wavelet transforms have either required non-uniform decimation -- leading to irregular data structures that are cumbersome to work with -- or require excessively high oversampling with unacceptable computational overhead. Here, we present a novel decimation strategy for wavelet transforms that leads to stable representations with oversampling rates close to one and uniform decimation. Specifically, we show that finite implementations of the resulting representation are energy-preserving in the sense of frame theory. The obtained wavelet coefficients can be stored in a timefrequency matrix with a natural interpretation of columns as time frames and rows as frequency channels. This matrix structure immediately grants access to a large number of algorithms that are successfully used in time-frequency audio processing, but could not previously be used jointly with wavelet transforms. We demonstrate the application of our method in processing based on nonnegative matrix factorization, in onset detection, and in phaseless reconstruction.

eess.AS↗

Efficient computation of the zeros of the Bargmann transform under additive white noise

We study the computation of the zero set of the Bargmann transform of a signal contaminated with complex white noise, or, equivalently, the computation of the zeros of its short-time Fourier transform with Gaussian window. We introduce the adaptive minimal grid neighbors algorithm (AMN), a variant of a method that has recently appeared in the signal processing literature, and prove that with high probability it computes the desired zero set. More precisely, given samples of the Bargmann transform of a signal on a finite grid with spacing $δ$, AMN is shown to compute the desired zero set up to a factor of $δ$ in the Wasserstein error metric, with failure probability $O(δ^4 \log^2(1/δ))$. We also provide numerical tests and comparison with other algorithms.

math.NA↗

A Differential Entropy Estimator for Training Neural Networks

Mutual Information (MI) has been widely used as a loss regularizer for training neural networks. This has been particularly effective when learn disentangled or compressed representations of high dimensional data. However, differential entropy (DE), another fundamental measure of information, has not found widespread use in neural network training. Although DE offers a potentially wider range of applications than MI, off-the-shelf DE estimators are either non differentiable, computationally intractable or fail to adapt to changes in the underlying distribution. These drawbacks prevent them from being used as regularizers in neural networks training. To address shortcomings in previously proposed estimators for DE, here we introduce KNIFE, a fully parameterized, differentiable kernel-based estimator of DE. The flexibility of our approach also allows us to construct KNIFE-based estimators for conditional (on either discrete or continuous variables) DE, as well as MI. We empirically validate our method on high-dimensional synthetic data and further apply it to guide the training of neural networks for real-world tasks. Our experiments on a large variety of tasks, including visual domain adaptation, textual fair classification, and textual fine-tuning demonstrate the effectiveness of KNIFE-based estimation. Code can be found at https://github.com/g-pichler/knife.

cs.LG↗

Zeros of Gaussian Weyl-Heisenberg functions and hyperuniformity of charge

We study Gaussian random functions on the complex plane whose stochastics are invariant under the Weyl-Heisenberg group (twisted stationarity). The theory is modeled on translation invariant Gaussian entire functions, but allows for non-analytic examples, in which case winding numbers can be either positive or negative. We calculate the first intensity of zero sets of such functions, both when considered as points on the plane, or as charges according to their phase winding. In the latter case, charges are shown to be in a certain average equilibrium independently of the particular covariance structure (universal screening). We investigate the corresponding fluctuations, and show that in many cases they are suppressed at large scales (hyperuniformity). This means that universal screening is empirically observable at large scales. We also derive an asymptotic expression for the charge variance. As a main application, we obtain statistics for the zero sets of the short-time Fourier transform of complex white noise with general windows, and also prove the following uncertainty principle: the expected number of zeros per unit area is minimized, among all window functions, exactly by generalized Gaussians. Further applications include poly-entire functions such as covariant derivatives of Gaussian entire functions.

math.PR↗

Fusion of Probability Density Functions

Fusing probabilistic information is a fundamental task in signal and data processing with relevance to many fields of technology and science. In this work, we investigate the fusion of multiple probability density functions (pdfs) of a continuous random variable or vector. Although the case of continuous random variables and the problem of pdf fusion frequently arise in multisensor signal processing, statistical inference, and machine learning, a universally accepted method for pdf fusion does not exist. The diversity of approaches, perspectives, and solutions related to pdf fusion motivates a unified presentation of the theory and methodology of the field. We discuss three different approaches to fusing pdfs. In the axiomatic approach, the fusion rule is defined indirectly by a set of properties (axioms). In the optimization approach, it is the result of minimizing an objective function that involves an information-theoretic divergence or a distance measure. In the supra-Bayesian approach, the fusion center interprets the pdfs to be fused as random observations. Our work is partly a survey, reviewing in a structured and coherent fashion many of the concepts and methods that have been developed in the literature. In addition, we present new results for each of the three approaches. Our original contributions include new fusion rules, axioms, and axiomatic and optimization-based characterizations; a new formulation of supra-Bayesian fusion in terms of finite-dimensional parametrizations; and a study of supra-Bayesian fusion of posterior pdfs for linear Gaussian models.

eess.SP↗

On the Estimation of Information Measures of Continuous Distributions

The estimation of information measures of continuous distributions based on samples is a fundamental problem in statistics and machine learning. In this paper, we analyze estimates of differential entropy in $K$-dimensional Euclidean space, computed from a finite number of samples, when the probability density function belongs to a predetermined convex family $\mathcal{P}$. First, estimating differential entropy to any accuracy is shown to be infeasible if the differential entropy of densities in $\mathcal{P}$ is unbounded, clearly showing the necessity of additional assumptions. Subsequently, we investigate sufficient conditions that enable confidence bounds for the estimation of differential entropy. In particular, we provide confidence bounds for simple histogram based estimation of differential entropy from a fixed number of samples, assuming that the probability density function is Lipschitz continuous with known Lipschitz constant and known, bounded support. Our focus is on differential entropy, but we provide examples that show that similar results hold for mutual information and relative entropy as well.

cs.IT↗

Modelling the Utility of Group Testing for Public Health Surveillance

In epidemic or pandemic situations, resources for testing the infection status of individuals may be scarce. Although group testing can help to significantly increase testing capabilities, the (repeated) testing of entire populations can exceed the resources of any country. We thus propose an extension of the theory of group testing that takes into account the fact that definitely specifying the infection status of each individual is impossible. Our theory builds on assigning to each individual an infection status (healthy/infected), as well as an associated cost function for erroneous assignments. This cost function is versatile, e.g., it could take into account that false negative assignments are worse than false positive assignments and that false assignments in critical areas, such as health care workers, are more severe than in the general population. Based on this model, we study the optimal use of a limited number of tests to minimize the expected cost. More specifically, we utilize information-theoretic methods to give a lower bound on the expected cost and describe simple strategies that can significantly reduce the expected cost over currently known strategies. A detailed example is provided to illustrate our theory.

q-bio.PE↗

Filtering with Wavelet Zeros and Gaussian Analytic Functions

We present the continuous wavelet transform (WT) of white Gaussian noise and establish a connection to the theory of Gaussian analytic functions. Based on this connection, we propose a methodology that detects components of a signal in white noise based on the distribution of the zeros of its continuous WT. To illustrate that the continuous theory can be employed in a discrete setting, we establish a uniform convergence result for the discretized continuous WT and apply the proposed method to a variety of acoustic signals.

math.NA↗

Minimal Achievable Sufficient Statistic Learning

We introduce Minimal Achievable Sufficient Statistic (MASS) Learning, a training method for machine learning models that attempts to produce minimal sufficient statistics with respect to a class of functions (e.g. deep networks) being optimized over. In deriving MASS Learning, we also introduce Conserved Differential Information (CDI), an information-theoretic quantity that - unlike standard mutual information - can be usefully applied to deterministically-dependent continuous random variables like the input and output of a deep network. In a series of experiments, we show that deep networks trained with MASS Learning achieve competitive performance on supervised learning and uncertainty quantification benchmarks.

cs.LG↗

Characterization of Analytic Wavelet Transforms and a New Phaseless Reconstruction Algorithm

We obtain a characterization of all wavelets leading to analytic wavelet transforms (WT). The characterization is obtained as a by-product of the theoretical foundations of a new method for wavelet phase reconstruction from magnitude-only coefficients. The cornerstone of our analysis is an expression of the partial derivatives of the continuous WT, which results in phase-magnitude relationships similar to the short-time Fourier transform (STFT) setting and valid for the generalized family of Cauchy wavelets. We show that the existence of such relations is equivalent to analyticity of the WT up to a multiplicative weight and a scaling of the mother wavelet. The implementation of the new phaseless reconstruction method is considered in detail and compared to previous methods. It is shown that the proposed method provides significant performance gains and a great flexibility regarding accuracy versus complexity. Additionally, we discuss the relation between scalogram reassignment operators and the wavelet transform phase gradient and present an observation on the phase around zeros of the WT.

math.NA↗

Rate-Distortion Theory of Finite Point Processes

We study the compression of data in the case where the useful information is contained in a set rather than a vector, i.e., the ordering of the data points is irrelevant and the number of data points is unknown. Our analysis is based on rate-distortion theory and the theory of finite point processes. We introduce fundamental information-theoretic concepts and quantities for point processes and present general lower and upper bounds on the rate-distortion function. To enable a comparison with the vector setting, we concretize our bounds for point processes of fixed cardinality. In particular, we analyze a fixed number of unordered Gaussian data points and show that we can significantly reduce the required rates compared to the best possible compression strategy for Gaussian vectors. As an example of point processes with variable cardinality, we study the best possible compression of Poisson point processes. For the specific case of a Poisson point process with uniform intensity on the unit square, our lower and upper bounds are separated by only a small gap and thus provide a good characterization of the rate-distortion function.

cs.IT↗