SearcharxivSearch

arXiv subjects

Kiril Ratmanski

Publications and source records attributed to Kiril Ratmanski.

3 recordsLinked to original sources

An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing

We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing. A lightweight neural controller predicts, in real time, the coefficients of a differentiable cascade of 35 second-order IIR filters (biquads), so the model tracks non-stationary noise while keeping a fully interpretable processing chain: every spectral modification is an explicit, adjustable equalizer curve rather than an opaque `black-box' transform. Because the biquad cascade carries the signal processing, the controller can be made very small, driving the cascade with only 24k parameters at a 10.7ms algorithmic latency, within hearing-aid budgets, and running entirely on-device so that audio never leaves the device. We also expose the suppression-versus-preservation trade-off as an explicit control: it can be set during training through the loss weighting, and adjusted at inference, with no retraining, by mixing the noisy input with the denoised output. On hearing-aid metrics (HASPI/HASQI) the 24k model stays within about 0.02 of DFNet3 (2.3M parameters, almost two orders of magnitude larger) while using about 29X fewer multiply-accumulates, although larger black-box models still lead on reference metrics such as PESQ. We present TVF as a proof of concept for a compact, interpretable, and controllable denoiser for on-device assistive hearing.

cs.SD

On real-time multi-stage speech enhancement systems

Recently, multi-stage systems have stood out among deep learning-based speech enhancement methods. However, these systems are always high in complexity, requiring millions of parameters and powerful computational resources, which limits their application for real-time processing in low-power devices. Besides, the contribution of various influencing factors to the success of multi-stage systems remains unclear, which presents challenges to reduce the size of these systems. In this paper, we extensively investigate a lightweight two-stage network with only 560k total parameters. It consists of a Mel-scale magnitude masking model in the first stage and a complex spectrum mapping model in the second stage. We first provide a consolidated view of the roles of gain power factor, post-filter, and training labels for the Mel-scale masking model. Then, we explore several training schemes for the two-stage network and provide some insights into the superiority of the two-stage network. We show that the proposed two-stage network trained by an optimal scheme achieves a performance similar to a four times larger open source model DeepFilterNet2.

eess.AS

Universally Rigid Framework Attachments

A framework is a graph and a map from its vertices to R^d. A framework is called universally rigid if there is no other framework with the same graph and edge lengths in R^d' for any d'. A framework attachment is a framework constructed by joining two frameworks on a subset of vertices. We consider an attachment of two universally rigid frameworks that are in general position in R^d. We show that the number of vertices in the overlap between the two frameworks must be sufficiently large in order for the attachment to remain universally rigid. Furthermore, it is shown that universal rigidity of such frameworks is preserved even after removing certain edges. Given positive semidefinite stress matrices for each of the two initial frameworks, we analytically derive the PSD stress matrices for the combined and edge-reduced frameworks. One of the benefits of the results is that they provide a general method for generating new universally rigid frameworks.

math.MG