arXiv · 2602.16399
Multi-Channel Replay Speech Detection using Acoustic Maps
Abstract
Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay speech detection from multi-channel recordings. Derived from classical beamforming over discrete azimuth and elevation grids, acoustic maps encode directional energy distributions that reflect physical differences between human speech radiation and loudspeaker-based replay. A lightweight convolutional neural network is designed to operate on this representation, achieving competitive performance on the ReMASC dataset with approximately 6k trainable parameters. Experimental results show that acoustic maps provide a compact and physically interpretable feature space for replay attack detection across different devices and acoustic environments.
Explore related subjects
Keep this discovery
Michael Neri, Tuomas Virtanen. 2026-02-18. Multi-Channel Replay Speech Detection using Acoustic Maps. https://arxiv.org/abs/2602.16399
Cite the original work for its findings. Save a collection to share your selection of sources.