SearcharxivSearch

arXiv subjects

Francesco Nesta

Publications and source records attributed to Francesco Nesta.

2 recordsLinked to original sources

FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses

This paper presents a novel multi-channel speech enhancement approach, FoVNet, that enables highly efficient speech enhancement within a configurable field of view (FoV) of a smart-glasses user without needing specific target-talker(s) directions. It advances over prior works by enhancing all speakers within any given FoV, with a hybrid signal processing and deep learning approach designed with high computational efficiency. The neural network component is designed with ultra-low computation (about 50 MMACS). A multi-channel Wiener filter and a post-processing module are further used to improve perceptual quality. We evaluate our algorithm with a microphone array on smart glasses, providing a configurable, efficient solution for augmented hearing on energy-constrained devices. FoVNet excels in both computational efficiency and speech quality across multiple scenarios, making it a promising solution for smart glasses applications.

cs.SD

Performance Analysis of Source Image Estimators in Blind Source Separation

Blind methods often separate or identify signals or signal subspaces up to an unknown scaling factor. Sometimes it is necessary to cope with the scaling ambiguity, which can be done through reconstructing signals as they are received by sensors, because scales of the sensor responses (images) have known physical interpretations. In this paper, we analyze two approaches that are widely used for computing the sensor responses, especially, in Frequency-Domain Independent Component Analysis. One approach is the least-squares projection, while the other one assumes a regular mixing matrix and computes its inverse. Both estimators are invariant to the unknown scaling. Although frequently used, their differences were not studied yet. A goal of this work is to fill this gap. The estimators are compared through a theoretical study, perturbation analysis and simulations. We point to the fact that the estimators are equivalent when the separated signal subspaces are orthogonal, and vice versa. Two applications are shown, one of which demonstrates a case where the estimators yield substantially different results.

cs.SD