arXiv · 2509.19495
ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement
Abstract
Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically study artifact prediction and reduction in diffusion-based SE. We show that variance in speech embeddings can be used to predict phonetic errors during inference. Building on these findings, we propose an ensemble inference method guided by semantic consistency across multiple diffusion runs. This technique reduces WER by 15% in low-SNR conditions, effectively improving phonetic accuracy and semantic plausibility. Finally, we analyze the effect of the number of diffusion steps, showing that adaptive diffusion steps balance artifact suppression and latency. Our findings highlight semantic priors as a powerful tool to guide generative SE toward artifact-free outputs.
Explore related subjects
Keep this discovery
Bhawana Chhaglani, Yang Gao, Julius Richter, Xilin Li, Syavosh Zadissa, Tarun Pruthi, Andrew Lovitt. 2025-09-23. ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement. https://arxiv.org/abs/2509.19495
Cite the original work for its findings. Save a collection to share your selection of sources.