Searcharxiv⌕ Search

arXiv · 2609.34524

Measurement-Based Bitrate-Energy-Quality Analysis of Neural Audio Codec Decoders on Laptop and Phone Platforms

Abstract

Neural audio codecs can achieve similar objective quality at lower bitrates than conventional codecs, but their decoder-side computational cost may offset this bitrate advantage on battery-powered client devices. This paper presents a measurement-based rate--energy--quality analysis of four neural audio codecs--EnCodec, DAC, HILCodec, and SNAC--and two conventional baselines, AAC-LC and Opus, on laptop and phone platforms. Speech and music are evaluated separately using the original-reference ViSQOL protocol. For the main comparison, operating points are matched by selecting the measured point nearest to the midpoint of the common overlap in treatment-level mean ViSQOL. A secondary analysis includes only explicitly evaluated bitrate settings. Decoder energy is measured as idle-subtracted energy per second of audio (J/s), and results are summarized as the median of three repeated runs for execution-valid runtime and device paths. Pairwise break-even transmission-energy thresholds are then derived analytically from the measured decoder energy and bitrate. At the matched-quality operating points, the evaluated neural codecs achieved similar ViSQOL scores at lower bitrates but generally required more decoder-side energy than the applicable conventional codecs. EnCodec produced the lowest neural break-even thresholds in both matched-quality cohorts on both platforms. By contrast, several DAC and SNAC comparisons on the Phone XNNPACK CPU path exceeded 200 mJ/kbit, and their full-band Phone configurations required more than 1 s to decode 1 s of audio. These results show that lower bitrate alone does not guarantee an energy benefit: deployment efficiency also depends on decoder complexity, the effective runtime and device mapping, execution validity, accelerator availability, and the transmission-energy coefficient.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Seunghyeon Shin, Seokjin Lee. 2026-09-28. Measurement-Based Bitrate-Energy-Quality Analysis of Neural Audio Codec Decoders on Laptop and Phone Platforms. https://arxiv.org/abs/2609.34524

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization

Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corresponding to objects heard in an audio signal. Most existing approaches tackle this problem by fine-tuning pre-trained models or by training additional modules specifically for the task. We adopt a different strategy: we introduce a training-free approach that leverages Non-negative Matrix Factorization (NMF) to co-factorize audio and visual features from pre-trained models so as to reveal shared interpretable concepts. These concepts are passed on to an open-vocabulary segmentation model for precise segmentation maps. By using frozen pre-trained models, our method achieves high generalization and establishes state-of-the-art performance in unsupervised sound-prompted segmentation, significantly surpassing previous unsupervised methods.

eess.AS↗

ImmersiveFlow: Stereo-to-7.1.4 spatial audio generation with flow matching

Immersive spatial audio has become increasingly critical for applications ranging from AR/VR to home entertainment and automotive sound systems. However, existing generative methods remain constrained to low-dimensional formats such as binaural audio and First-Order Ambisonics (FOA). Binaural rendering is inherently limited to headphone playback, while FOA suffers from spatial aliasing and insufficient resolution for high-frequency. To overcome these limitations, we introduce ImmersiveFlow, the first end-to-end generative framework that directly synthesizes discrete 7.1.4 format spatial audio from stereo input. ImmersiveFlow leverages Flow Matching to learn trajectories from stereo inputs to multichannel spatial features within a pretrained VAE latent space. At inference, the Flow Matching model predicted latent features are decoded by the VAE and converted into the final 7.1.4 waveform. Comprehensive objective and subjective evaluations demonstrate that our method produces perceptually rich sound fields and enhanced spatial impression, significantly outperforming traditional upmixing techniques.

eess.AS↗

Data Augmentation for L2 English Speaking Assessment using TTS

Automated assessment of second language (L2) speaking proficiency requires substantial annotated speech data, which are scarce compared to written learner corpora. We investigate whether written L2 responses can be transformed into useful synthetic speech for proficiency assessment using text-to-speech (TTS) and voice cloning. Using COREFL, a corpus of paired spoken and written responses from L2 learners of English, we systematically study two factors: how written responses should be transformed into spoken-style language ("speechification") and how synthetic voices and texts should be paired based on shared learner attributes (proficiency level, first language, both, or neither). We generate speechified responses with a large language model and synthesise them using TTS and voice cloning, then evaluate their utility for audio-based (HuBERT) and text-based (ModernBERT) proficiency grading. Results show that pairing voices and texts based on proficiency provides a principled strategy for synthetic data generation, while speechification substantially improves the match between written and spoken L2 and improves downstream grading performance. Augmenting real training data with synthetic speechified responses improves both HuBERT- and ModernBERT-based graders, demonstrating the potential of synthetic spoken data for L2 speaking proficiency assessment.

eess.AS↗