SearcharxivSearch

arXiv subjects

Renzheng Shi

Publications and source records attributed to Renzheng Shi.

2 recordsLinked to original sources

EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase

The Griffin-Lim algorithm has been a seminal contribution for phase reconstruction from amplitude spectrograms, however, requiring (infinitely) high algorithmic delay. Its low-delay variant suffers in speech quality. Recent (generative) neural network methods improve on speech quality still at medium to high algorithmic delay, but often they are complex and optimized only for one specific input representation. We build upon an efficient low-delay speech vocoder and propose EffVOC, which supports synthesis of wideband (WB) or fullband (FB) speech from either amplitude spectrum or Mel coefficient inputs. We evaluate both input representations across multiple model sizes in a unified framework and compare against state of the art. Results show that our proposed low-delay (20 ms vs. 32 ms or more) efficient approach marks a new SOTA by achieving top-ranked subjective MOS scores (WB: 4.17/4.15, FB: 4.14/4.11) for amplitude spectrum/Mel representations, very close to ground truth.

eess.AS

Non-Causal to Causal SSL-Supported Transfer Learning: Towards a High-Performance Low-Latency Speech Vocoder

Recently, BigVGAN has emerged as high-performance speech vocoder. Its sequence-to-sequence-based synthesis, however, prohibits usage in low-latency conversational applications. Our work addresses this shortcoming in three steps. First, we introduce low latency into BigVGAN via implementing causal convolutions, yielding decreased performance. Second, to regain performance, we propose a teacher-student transfer learning scheme to distill the high-delay non-causal BigVGAN into our low-latency causal vocoder. Third, taking advantage of a self-supervised learning (SSL) model, in our case wav2vec 2.0, we align its encoder speech representations extracted from our low-latency causal vocoder to the ground truth ones. In speaker-independent settings, both proposed training schemes notably elevate the performance of our low-latency vocoder, closing up to the original high-delay BigVGAN. At only 21% higher complexity, our best small causal vocoder achieves 3.96 PESQ and 1.25 MCD, excelling even the original small non-causal BigVGAN (3.64 PESQ) by 0.32 PESQ and 0.1 MCD points, respectively.

eess.AS