Searcharxiv⌕ Search

arXiv · 2609.36379

UDSS-BWE: Uncertainty- and Decision-Science Inspired Swin BandWidth Extension

Abstract

Bandwidth extension (BWE) is fundamentally localized: the most perceptual distortions are not average-case distortions, but rare high-frequency (HF) transients that standard, risk-neutral objectives tend to smooth away. To close this gap, we seek solutions in the risk-sensitive and uncertainty-aware decision science rules and present UDSS-BWE, which introduces five decision-science and uncertainty-aware discriminators: CVaRD (does tail pooling to amplify HF artifacts), CCD (a primal-dual augmented Lagrangian to prevent HF overboost), MCUD (a learnable utility over spectral flatness/ centroid/ rolloff), EDD (captures epistemic uncertainty), and DROD (captures entropic KL-DRO aggregation). UDSS-BWE is also designed as a complex valued adversarial BWE framework that uses Swin-based generators, a lightweight dual-stream shifted-window backbone, to capture local and long-range structure efficiently, while learnable lattice coupling provides controlled cross-stream exchange. UDSS-BWE is optimized extensively and achieves better perceptual quality with 3.89x fewer parameters (72M vs.18.5M) over two English and French datasets under clean and noisy conditions. To the best of our knowledge, this work shows how multi disciplinary decision-science-inspired and uncertainty theories can be successfully used to design efficient discriminators for producing more nuanced audios, establishing a new baseline in the BWE task.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tarikul Islam Tamiti, Sajid Fardin Dipto, David Vergano, Luke Baja-Ricketts, Anomadarshi Barua. 2026-09-28. UDSS-BWE: Uncertainty- and Decision-Science Inspired Swin BandWidth Extension. https://arxiv.org/abs/2609.36379

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension

Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.

cs.SD↗

InterBias-SV: Compound Conditions in Speaker Verification

Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.

cs.SD↗

Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code

cs.SD↗