SearcharxivSearch

arXiv subjects

Benjamin Stahl

Publications and source records attributed to Benjamin Stahl.

6 recordsLinked to original sources

Residual Learning for Neural Ambisonics Encoders

Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provides a device-agnostic spatial audio representation by mapping array signals to spherical harmonic (SH) coefficients. In practice, however, accurate encoding remains challenging. While traditional linear encoders are signal-independent and robust, they amplify low-frequency noise and suffer from high-frequency spatial aliasing. On the other hand, neural network approaches can outperform linear encoders but they often assume idealized microphones and may perform inconsistently in real-world scenarios. To leverage their complementary strengths, we introduce a residual-learning framework that refines a linear encoder with corrections from a neural network. Using measured array transfer functions from smartglasses, we compare a UNet-based encoder from the literature with a new recurrent attention model. Our analysis reveals that both neural encoders only consistently outperform the linear baseline when integrated within the residual learning framework. In the residual configuration, both neural models achieve consistent and significant improvements across all tested metrics for in-domain data and moderate gains for out-of-domain data. Yet, coherence analysis indicates that all neural encoder configurations continue to struggle with directionally accurate high-frequency encoding.

eess.AS

Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce nonlinear relationships between the separated and reference signals, limiting the reliability of these metrics for objective evaluation. To address this issue, we conduct a Degradation Category Rating listening test and analyze correlations between the obtained degradation mean opinion scores (DMOS) and a set of objective audio quality metrics for the task of singing voice separation. We evaluate three state-of-the-art discriminative models and two new competitive generative models. For both discriminative and generative models, intrusive embedding-based metrics show higher correlations with DMOS than conventional intrusive metrics such as BSS-Eval. For discriminative models, the highest correlation is achieved by the MSE computed on Music2Latent embeddings. When it comes to the evaluation of generative models, the strongest correlations are evident for the multi-resolution STFT loss and the MSE calculated on MERT-L12 embeddings, with the latter also providing the most balanced correlation across both model types. Our results highlight the limitations of BSS-Eval metrics for evaluating generative singing voice separation models and emphasize the need for careful selection and validation of alternative evaluation metrics for the task of singing voice separation.

eess.AS

Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment

In this paper, we investigate distillation and pruning methods to reduce model size for non-intrusive speech quality assessment based on self-supervised representations. Our experiments build on XLS-R-SQA, a speech quality assessment model using wav2vec 2.0 XLS-R embeddings. We retrain this model on a large compilation of mean opinion score datasets, encompassing over 100,000 labeled clips. For distillation, using this model as a teacher, we generate pseudo-labels on unlabeled degraded speech signals and train student models of varying sizes. For pruning, we use a data-driven strategy. While data-driven pruning performs better at larger model sizes, distillation on unlabeled data is more effective for smaller model sizes. Distillation can halve the gap between the baseline's correlation with ground-truth MOS labels and that of the XLS-R-based teacher model, while reducing model size by two orders of magnitude compared to the teacher model.

eess.AS

SN 2016esw: a luminous Type II supernova observed within the first day after the explosion

We present photometry, spectroscopy, and host-galaxy integral-field spectroscopy of the Type II supernova (SN) 2016esw in CGCG~229-009 from the first day after the explosion up to 120 days. Its light-curve shape is similar to that of a typical SN II; however, SN 2016esw is near the high-luminosity end of the SN II distribution, with a peak of $M^{\rm max}_{V}=-18.36$ mag. The $V$-band light curve exhibits a long recombination phase for a SN II (similar to the long-lived plateau of SN 2004et). Considering the well-known relation between the luminosity and the plateau decline rate, SN 2016esw should have a $V$-band slope of $\sim 2.10$ mag (100 days)$^{-1}$; however, SN 2016esw has a substantially flatter plateau with a slope of $1.01\pm 0.26$ mag (100 days)$^{-1}$, perhaps indicating that interacting Type II supernovae are not useful for cosmology. At 19.5 days post-explosion, the spectrum presents a boxy H$α$ emission line with flat absorption profiles, suggesting interaction between the ejecta and circumstellar matter. Finally, based on the spectral properties, SN 2016esw shows similarities with the luminous and interacting SN 2007pk at early epochs, particularly in terms of observable line features and their evolution.

astro-ph.HE

Discovery and Follow-up Observations of the Young Type Ia Supernova 2016coj

The Type~Ia supernova (SN~Ia) 2016coj in NGC 4125 (redshift $z=0.004523$) was discovered by the Lick Observatory Supernova Search 4.9 days after the fitted first-light time (FFLT; 11.1 days before $B$-band maximum). Our first detection (pre-discovery) is merely $0.6\pm0.5$ day after the FFLT, making SN 2016coj one of the earliest known detections of a SN Ia. A spectrum was taken only 3.7 hr after discovery (5.0 days after the FFLT) and classified as a normal SN Ia. We performed high-quality photometry, low- and high-resolution spectroscopy, and spectropolarimetry, finding that SN 2016coj is a spectroscopically normal SN Ia, but with a high velocity of \ion{Si}{2} $λ$6355 ($\sim 12,600$\,\kms\ around peak brightness). The \ion{Si}{2} $λ$6355 velocity evolution can be well fit by a broken-power-law function for up to a month after the FFLT. SN 2016coj has a normal peak luminosity ($M_B \approx -18.9 \pm 0.2$ mag), and it reaches a $B$-band maximum \about16.0~d after the FFLT. We estimate there to be low host-galaxy extinction based on the absence of Na~I~D absorption lines in our low- and high-resolution spectra. The spectropolarimetric data exhibit weak polarization in the continuum, but the \ion{Si}{2} line polarization is quite strong ($\sim 0.9\% \pm 0.1\%$) at peak brightness.

astro-ph.SR

The XMM Cluster Survey: evolution of the velocity dispersion -- temperature relation over half a Hubble time

We measure the evolution of the velocity dispersion--temperature ($σ_{\rm v}$--$T_{\rm X}$) relation up to $z = 1$ using a sample of 38 galaxy clusters drawn from the \textit{XMM} Cluster Survey. This work improves upon previous studies by the use of a homogeneous cluster sample and in terms of the number of high redshift clusters included. We present here new redshift and velocity dispersion measurements for 12 $z > 0.5$ clusters observed with the GMOS instruments on the Gemini telescopes. Using an orthogonal regression method, we find that the slope of the relation is steeper than that expected if clusters were self-similar, and that the evolution of the normalisation is slightly negative, but not significantly different from zero ($σ_{\rm v} \propto T^{0.86 \pm 0.14} E(z)^{-0.37 \pm 0.33}$). We verify our results by applying our methods to cosmological hydrodynamical simulations. The lack of evolution seen in our data is consistent with simulations that include both feedback and radiative cooling.

astro-ph.CO