SearcharxivSearch

arXiv subjects

Aleksandr Zaitsev

Publications and source records attributed to Aleksandr Zaitsev.

2 recordsLinked to original sources

Observation of $Π$-symmetry ultralong-range Rydberg molecules

We observe weakly-bound $Π$-symmetry electronic states in the spectroscopy of $^{87}$Rb$(nP_{3/2})$+$^{87}$Rb($5S_{1/2}$) ultralong-range Rydberg molecules. We detect these molecules in Rydberg states having principal quantum number $13\le n \le 16$. Their $Π$-state character is unambiguously identified via their observed multiplet structure: the $2F+1$ magnetic sublevels of the ground-state rubidium atom separate, as in the Zeeman effect, because of the spin-spin coupling between the Rydberg and valence electrons. We find a rapid decrease in the molecular binding energy $\propto (n-μ_{P_{3/2}})^{-11}$, where $μ_{P_{3/2}}$ is the quantum defect, indicating that the low-$n$ regime of Rydberg states is ideally suited for studies of $Π$-symmetry molecules. Our observations are in good agreement with Green's function-based calculations for $14\le n\le 16$, with poorer agreement for $n=13$ hinting at the beginning of a breakdown of the Fermi pseudopotential approach at low $n$.

physics.atom-ph

Neural networks for Text-to-Speech evaluation

Ensuring that Text-to-Speech (TTS) systems deliver human-perceived quality at scale is a central challenge for modern speech technologies. Human subjective evaluation protocols such as Mean Opinion Score (MOS) and Side-by-Side (SBS) comparisons remain the de facto gold standards, yet they are expensive, slow, and sensitive to pervasive assessor biases. This study addresses these barriers by formulating, and implementing a suite of novel neural models designed to approximate expert judgments in both relative (SBS) and absolute (MOS) settings. For relative assessment, we propose NeuralSBS, a HuBERT-backed model achieving 73.7% accuracy (on SOMOS dataset). For absolute assessment, we introduce enhancements to MOSNet using custom sequence-length batching, as well as WhisperBert, a multimodal stacking ensemble that combines Whisper audio features and BERT textual embeddings via weak learners. Our best MOS models achieve a Root Mean Square Error (RMSE) of ~0.40, significantly outperforming the human inter-rater RMSE baseline of 0.62. Furthermore, our ablation studies reveal that naively fusing text via cross-attention can degrade performance, highlighting the effectiveness of ensemble-based stacking over direct latent fusion. We additionally report negative results with SpeechLM-based architectures and zero-shot LLM evaluators (Qwen2-Audio, Gemini 2.5 flash preview), reinforcing the necessity of dedicated metric learning frameworks.

cs.CL