SearcharxivSearch

arXiv subjects

Jan Büthe

Publications and source records attributed to Jan Büthe.

13 recordsLinked to original sources

A lightweight and robust method for blind wideband-to-fullband extension of speech

Reducing the bandwidth of speech is common practice in resource constrained environments like low-bandwidth speech transmission or low-complexity vocoding. We propose a lightweight and robust method for extending the bandwidth of wideband speech signals that is inspired by classical methods developed in the speech coding context. The resulting model has just ~370K parameters and a complexity of ~140 MFLOPS (or ~70 MMACS). With a frame size of 10 ms and a lookahead of only 0.27 ms, the model is well-suited for use with common wideband speech codecs. We evaluate the model's robustness by pairing it with the Opus SILK speech codec (1.5 release) and verify in a P.808 DCR listening test that it significantly improves quality from 6 to 12 kb/s. We also demonstrate that Opus 1.5 together with the proposed bandwidth extension at 9 kb/s meets the quality of 3GPP EVS at 9.6 kb/s and that of Opus 1.4 at 18 kb/s showing that the blind bandwidth extension can meet the quality of classical guided bandwidth extensions thus providing a way for backward-compatible quality improvement.

eess.AS

DRED: Deep REDundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder

Despite recent advancements in packet loss concealment (PLC) using deep learning techniques, packet loss remains a significant challenge in real-time speech communication. Redundancy has been used in the past to recover the missing information during losses. However, conventional redundancy techniques are limited in the maximum loss duration they can cover and are often unsuitable for burst packet loss. We propose a new approach based on a rate-distortion-optimized variational autoencoder (RDO-VAE), allowing us to optimize a deep speech compression algorithm for the task of encoding large amounts of redundancy at very low bitrate. The proposed Deep REDundancy (DRED) algorithm can transmit up to 50x redundancy using less than 32 kb/s. Results show that DRED outperforms the existing Opus codec redundancy. We also demonstrate its benefits when operating in the context of WebRTC.

eess.AS

Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction

Neural vocoders are now being used in a wide range of speech processing applications. In many of those applications, the vocoder can be the most complex component, so finding lower complexity algorithms can lead to significant practical benefits. In this work, we propose FARGAN, an autoregressive vocoder that takes advantage of long-term pitch prediction to synthesize high-quality speech in small subframes, without the need for teacher-forcing. Experimental results show that the proposed 600~MFLOPS FARGAN vocoder can achieve both higher quality and lower complexity than existing low-complexity vocoders. The quality even matches that of existing higher-complexity vocoders.

eess.AS

NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping

Speech codec enhancement methods are designed to remove distortions added by speech codecs. While classical methods are very low in complexity and add zero delay, their effectiveness is rather limited. Compared to that, DNN-based methods deliver higher quality but they are typically high in complexity and/or require delay. The recently proposed Linear Adaptive Coding Enhancer (LACE) addresses this problem by combining DNNs with classical long-term/short-term postfiltering resulting in a causal low-complexity model. A short-coming of the LACE model is, however, that quality quickly saturates when the model size is scaled up. To mitigate this problem, we propose a novel adatpive temporal shaping module that adds high temporal resolution to the LACE model resulting in the Non-Linear Adaptive Coding Enhancer (NoLACE). We adapt NoLACE to enhance the Opus codec and show that NoLACE significantly outperforms both the Opus baseline and an enlarged LACE model at 6, 9 and 12 kb/s. We also show that LACE and NoLACE are well-behaved when used with an ASR system.

eess.AS

LACE: A light-weight, causal model for enhancing coded speech through adaptive convolutions

Classical speech coding uses low-complexity postfilters with zero lookahead to enhance the quality of coded speech, but their effectiveness is limited by their simplicity. Deep Neural Networks (DNNs) can be much more effective, but require high complexity and model size, or added delay. We propose a DNN model that generates classical filter kernels on a per-frame basis with a model of just 300~K parameters and 100~MFLOPS complexity, which is a practical complexity for desktop or mobile device CPUs. The lack of added delay allows it to be integrated into the Opus codec, and we demonstrate that it enables effective wideband encoding for bitrates down to 6 kb/s.

eess.AS

Framewise WaveGAN: High Speed Adversarial Vocoder in Time Domain with Very Low Computational Complexity

GAN vocoders are currently one of the state-of-the-art methods for building high-quality neural waveform generative models. However, most of their architectures require dozens of billion floating-point operations per second (GFLOPS) to generate speech waveforms in samplewise manner. This makes GAN vocoders still challenging to run on normal CPUs without accelerators or parallel computers. In this work, we propose a new architecture for GAN vocoders that mainly depends on recurrent and fully-connected networks to directly generate the time domain signal in framewise manner. This results in considerable reduction of the computational cost and enables very fast generation on both GPUs and low-complexity CPUs. Experimental results show that our Framewise WaveGAN vocoder achieves significantly higher quality than auto-regressive maximum-likelihood vocoders such as LPCNet at a very low complexity of 1.2 GFLOPS. This makes GAN vocoders more practical on edge and low-power devices.

eess.AS

Estimating $π(x)$ and related functions under partial RH assumptions

The aim of this paper is to give a direct interpretation of the validity of the Riemann hypothesis up to a certain height $T$ in terms of the prime-counting function $π(x)$. This is done by proving the well-known explicit Schoenfeld bound on the RH to hold as long as $4.92 \sqrt{x/\log(x)} \leq T$. Similar statements are proven for the Riemann prime-counting function and the Chebyshov functions $ψ(x)$ and $\vartheta(x)$. Apart from that, we also improve some of the existing bounds of Chebyshov type for the function $ψ(x)$.

math.NT

A Streamwise GAN Vocoder for Wideband Speech Coding at Very Low Bit Rate

Recently, GAN vocoders have seen rapid progress in speech synthesis, starting to outperform autoregressive models in perceptual quality with much higher generation speed. However, autoregressive vocoders are still the common choice for neural generation of speech signals coded at very low bit rates. In this paper, we present a GAN vocoder which is able to generate wideband speech waveforms from parameters coded at 1.6 kbit/s. The proposed model is a modified version of the StyleMelGAN vocoder that can run in frame-by-frame manner, making it suitable for streaming applications. The experimental results show that the proposed model significantly outperforms prior autoregressive vocoders like LPCNet for very low bit rate speech coding, with computational complexity of about 5 GMACs, providing a new state of the art in this domain. Moreover, this streamwise adversarial vocoder delivers quality competitive to advanced speech codecs such as EVS at 5.9 kbit/s on clean speech, which motivates further usage of feed-forward fully-convolutional models for low bit rate speech coding.

eess.AS

An analytic method for bounding $\psi(x)$

In this paper we present an analytic altorithm which calculates almost sharp bounds for the normalized error term $(t-\psi(t))/\sqrt{t}$ for $t\leq x$ in expected run time $O(x^{1/2+\varepsilon})$ for every $\varepsilon>0$. The method has been implemented and used to calculate the bound $|\psi(t) - t| \leq 0.94 \sqrt{t}$ for $11< t\leq 10^{19}$. In particular, this bound implies that $\operatorname{li}(t) - \pi(t) > 0$ for $t\in [2,10^{19}]$, which gives an improved lower bound for the Skewes number.

math.NT

On the first sign change in Mertens' Theorem

The function $\sum_{p\leq x} \frac{1}{p} - \log\log(x) - M$ is known to change sign infinitely often, but so far all calculated values are positive. In this paper we prove that the first sign change occurs well before $\exp(495.702833165)$.

math.NT

A method for proving the completeness of a list of zeros of certain L-functions

When it comes to partial numerical verification of the Riemann Hypothesis, one crucial part is to verify the completeness of a list of pre-computed zeros. Turing developed such a method, based on an explicit version of a theorem of Littlewood on the average of the argument of the Riemann zeta function. In a previous paper we suggested an alternative method based on the Weil-Barner explicit formula. This method asymptotically sacrifices fewer zeros in order to prove the completeness of a list of zeros with imaginary part in a given interval. In this paper, we prove a general version of this method for an extension of the Selberg class including Hecke and Artin L-series, L-functions of modular forms, and, at least in the unramified case, automorphic L-functions. As an example, we further specify this method for Hecke L-series and L-functions of elliptic curves over the rational numbers.

math.NT

An improved analytic Method for calculating $\pi(x)$

We present an improved version of the analytic method for calculating $\pi(x)$, the number of prime numbers not exceeding $x$. We implemented this method in cooperation with J. Franke, T. Kleinjung and A. Jost and calculated the value $\pi(10^{25})$.

math.NT