SearcharxivSearch

arXiv subjects

Siqi Zheng

Publications and source records attributed to Siqi Zheng.

At least 19 recordsLinked to original sources

NO molecule in massive star forming regions

Context. Among diatomic molecules composed of the abundant elements C, N and O, NO has been detected far less than the well studied CN and CO, making it a crucial yet under-observed component in nitrogen-containing chemical networks. NO was thought to serve as a potential tracer of shocks, as evidenced with orders abundance enhancements reported in literature. Aims. Large-sample observations for NO molecule in widespread interstellar environments are needed to confirm if the enhancement of NO is due to shock chemistry or not. Methods. Single-point survey for NO lines around 150 GHz was carried out by Arizona Radio Observatory 12-meter telescope towards a sample of 36 massive star forming regions containing SiO emission, which include three evolutionary stages: 4 IRDCs, 6 protostars and 26 H II regions. Results. The NO emission was detected in 28 sources with a detection rate of 78%. Beam-averaged NO column densities and abundances relative to H2 were derived from integrated intensities of two main hyperfine lines. Correlations between NO and SiO in integrated intensity and relative abundance are similar to the corresponding correlations of c-C3H2, indicating that NO enrichment may not significantly involve pronounced shock activities, which coincides with the trend in line widths: NO is close to c-C3H2, both smaller than H2CO and far smaller than SiO. Conclusions. Observational evidence does not strongly support significant NO enhancement by shock chemistry in the observed sources, indicating that the formation of NO does not necessarily require shocks.

astro-ph.GA

X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs

While the shift from cascaded dialogue systems to end-to-end (E2E) speech Large Language Models (LLMs) improves latency and paralinguistic modeling, E2E models often exhibit a significant performance degradation compared to their text-based counterparts. The standard Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) training methods fail to close this gap. To address this, we propose X-OPD, a novel Cross-Modal On-Policy Distillation framework designed to systematically align the capabilities of Speech LLMs to their text-based counterparts. X-OPD enables the Speech LLM to explore its own distribution via on-policy rollouts, where a text-based teacher model evaluates these trajectories and provides token-level feedback, effectively distilling teacher's capabilities into student's multi-modal representations. Extensive experiments across multiple benchmarks demonstrate that X-OPD significantly narrows the gap in complex tasks while preserving the model's inherent capabilities.

eess.AS

Confidence, Statistical Evidence and Relative Belief with Applications to a Problem in Particle Physics

Probability theory provides a clear definition of what is meant by evidence in favor, against or none either way, of an event occurring for an unobserved response, via the principle of evidence. This is immediately applicable when carrying out a proper Bayesian analysis. Even without a prior, this imposes restrictions on reported inferences as these need to reflect the likelihood ordering. Relative belief inferences satisfy this requirement and, when the errors in these inferences are controlled, they also satisfy repeated sampling, or frequentist, requirements such as achieving given confidence levels. Relative belief inferences are considered here for the construction of intervals for uncertainty quantification in the context of a Poisson model for a signal with background noise. These intervals are contrasted with the well-known Feldman-Cousins intervals for this problem.

physics.data-an

AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features

With the rapid advancement of smart glasses, voice interaction has been widely adopted due to its naturalness and convenience. However, its practical deployment is often undermined by vulnerability to spoofing attacks, while no public dataset currently exists for voice liveness detection and authentication in smart-glasses scenarios. To address this challenge, we first collect a multi-acoustic-modal dataset comprising 16-channel audio data from 42 subjects, along with corresponding attack samples covering two attack categories. Based on insights derived from this collected data, we propose AuthG-Live, a sound-field-based voice liveness detection method, and AuthG-Net, a multi-acoustic-modal authentication model. We further benchmark seven voice liveness detection methods and four authentication methods across diverse acoustic modalities. The results demonstrate that our proposed approach achieves state-of-the-art performance on four benchmark tasks, and extensive ablation studies validate the generalizability of our methods \red{under real-world constraints}. Finally, we release this dataset, termed AuthGlass, to facilitate future research on voice liveness detection and authentication for smart glasses.

cs.HC

The evolution of C4H and c-C3H2 in molecular cores

Linear C4H and cyclic c-C3H2, as small unsaturated hydrocarbons, are the key precursors to complex organic molecules and are critical components of the interstellar medium. We present on-the-fly mapping observations of C4H 9-8 lines, c-C3H2 2-1, H13CO+ 1-0, and H42 toward a sample of 22 massive star-forming regions using the IRAM 30m telescope. Our aim is to further explore the evolution of these carbon-chain molecules by combining observational results obtained in cold cores. We employed H13CO+ 1-0 and H42 as tracers to probe the positions of molecular cloud cores and ionised hydrogen regions (HII regions), respectively. One chemical model in particular, which includes gas, dust grain surface, and icy mantle phases for C4H and c-C3H2 molecules, was used to make comparisons with observed abundances. From mapping observations targeting 31 regions across 22 sources, C4H 9-8 (J = 19/2-17/2) and C4H 9-8 (J = 17/2-15/2) were detected in only 17 regions, while H13CO+ 1-0 and c-C3H2 2-1 were successfully detected in all 31 regions. We find that the emission of C4H 9-8 and c-C3H2 2-1 is concentrated at the edges of H42 emission regions. The C4H/H13CO+ and c-C3H2/H13CO+ relative abundance ratios range from 0.17 to 1.77 and 1.42 to 6.69, respectively, with a median C4H/c-C3H2 ratio of 0.13. By combining the observational results of cold cores, we find that C4H/H13CO+ and c-C3H2/H13CO+ ratios show a strong decreasing trend as molecular cores evolve. The decreasing trends in C4H/H13CO+ and c-C3H2/H13CO+ ratios imply that small unsaturated hydrocarbons can be consumed and converted into other organic molecules during the evolution of molecular cores. The spatial concentration of C4H and c-C3H2 emission at the edges of H42 regions further supports their role as precursors in the chemical pathways that lead to complex organic molecules in the interstellar medium.

astro-ph.GA

CH3CCH as a thermometer in warm molecular gas

Kinetic temperature is a fundamental parameter in molecular clouds. Symmetric top molecules, such as NH$_3$ and CH$_3$CCH, are often used as thermometers. However, at high temperatures, NH$_3$(2,2) can be collisionally excited to NH$_3$(2,1) and rapidly decay to NH$_3$(1,1), which can lead to an underestimation of the kinetic temperature when using rotation temperatures derived from NH$_3$(1,1) and NH$_3$(2,2). In contrast, CH$_3$CCH is a symmetric top molecule with lower critical densities of its rotational levels than those of NH$_3$, which can be thermalized close to the kinetic temperature at relatively low densities of about 10$^{4}$ cm$^{-3}$. To compare the rotation temperatures derived from NH$_3$(1,1)$\&$(2,2) and CH$_3$CCH rotational levels in warm molecular gas, we used observational data toward 55 massive star-forming regions obtained with Yebes 40m and TMRT 65m. Our results show that rotation temperatures derived from NH$_3$(1,1)$\&$(2,2) are systematically lower than those from CH$_3$CCH 5-4. This suggests that CH$_3$CCH rotational lines with the same $J$+1$\rightarrow$$J$ quantum number may be a more reliable thermometer than NH$_3$(1,1)$\&$(2,2) in warm molecular gas located in the surroundings of massive young stellar objects or, more generally, in massive star-forming regions.

astro-ph.GA

Widespread Hot Molecular Gas Heated by Shear-induced Turbulence in the Galactic Center

We observed NH3 metastable inversion lines from (3, 3) to (18, 18) toward G0.66-0.13 in the Galactic center with the Shanghai Tianma 65m radio telescope and Yebes 40 m telescope. Highly-excited lines of NH3 (17, 17), (18, 18) were detected in emission for the first time in the interstellar medium, with upper energy levels up to 3100 K. Mapping observations reveal widespread hot molecular gas traced by NH3 (13, 13) toward G0.66-0.13. The rotation temperatures of hot gas traced by NH3 exceed 400 K, which amounts to five percent of the total NH3 in the Galactic Center. Hot gas (>400 K) and warm gas (100-140 K) are found in distinct clumps, with the hot gas located at the interfacing regions between different warm clouds. The theory of intermittency in turbulence reproduces the complex temperature structure in the central molecular zone, especially the hot gas observed here. The results presented here demonstrate that turbulence heating dominates the heating of the molecular gas in the Central Molecular Zone, while the turbulence is induced by the shear-motion of molecular clouds under the gravitational potential of the nuclear star clusters and the supermassive black hole. Our results suggest that shear-induced turbulence heating could be a widespread factor influencing galactic evolution.

astro-ph.GA

WavReward: Spoken Dialogue Models With Generalist Reward Evaluators

End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' conversational performance has largely been overlooked. This is primarily due to the intelligent chatbots convey a wealth of non-textual information which cannot be easily measured using text-based language models like ChatGPT. To address this gap, we propose WavReward, a reward feedback model based on audio language models that can evaluate both the IQ and EQ of spoken dialogue systems with speech input. Specifically, 1) based on audio language models, WavReward incorporates the deep reasoning process and the nonlinear reward mechanism for post-training. By utilizing multi-sample feedback via the reinforcement learning algorithm, we construct a specialized evaluator tailored to spoken dialogue models. 2) We introduce ChatReward-30K, a preference dataset used to train WavReward. ChatReward-30K includes both comprehension and generation aspects of spoken dialogue models. These scenarios span various tasks, such as text-based chats, nine acoustic attributes of instruction chats, and implicit chats. WavReward outperforms previous state-of-the-art evaluation models across multiple spoken dialogue scenarios, achieving a substantial improvement about Qwen2.5-Omni in objective accuracy from 53.4$\%$ to 91.5$\%$. In subjective A/B testing, WavReward also leads by a margin of 83$\%$. Comprehensive ablation studies confirm the necessity of each component of WavReward. All data and code will be publicly at https://github.com/jishengpeng/WavReward after the paper is accepted.

eess.AS

ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker's voice without further control and adjustment capabilities while prior controllable TTS models cannot perform speaker-specific voice generation. Therefore, ControlSpeech focuses on a more challenging task: a TTS system with controllable timbre, content, and style at the same time. ControlSpeech takes speech prompts, content prompts, and style prompts as inputs and utilizes bidirectional attention and mask-based parallel decoding to capture codec representations corresponding to timbre, content, and style in a discrete decoupling codec space. Moreover, we analyze the many-to-many issue in textual style control and propose the Style Mixture Semantic Density (SMSD) module, which is based on Gaussian mixture density networks, to resolve this problem. To facilitate empirical validations, we make available a new style controllable dataset called VccmDataset. Our experimental results demonstrate that ControlSpeech exhibits comparable or state-of-the-art (SOTA) performance in terms of controllability, timbre similarity, audio quality, robustness, and generalizability. The relevant code and demo are available at https://github.com/jishengpeng/ControlSpeech .

eess.AS

Dual-band Unified Exploration of Three CMZ Clouds (DUET). Cloud-wide census of continuum sources showing low spectral indices

The Milky Way's Central Molecular Zone (CMZ) is measured to form stars 10 times less efficiently than in the Galactic disk, based on emission from high-mass stars. However, the CMZ's low-mass protostellar population, which accounts for most of the initial stellar mass budget and star formation rate (SFR), is poorly constrained observationally due to limited sensitivity and resolution. We present the Dual-band Unified Exploration of Three CMZ Clouds (DUET) survey, targeting the 20 km/s Cloud, Sgr C, and Dust Ridge cloud e using the Atacama Large Millimeter/submillimeter Array (ALMA) at 1.3 and 3 mm. The mosaicked observations achieve a comparable resolution of 0.2-0.3" (~1600-2500 au) and a sky coverage of 8.3-10.4 square arcmin, respectively. We report 563 continuum sources at 1.3 mm and 330 at 3 mm, respectively, and a dual-band catalog with 450 continuum sources. These sources are marginally resolved at the 2,000 au resolution. We find a cloud-wide deviation (>70%) from commonly-used dust modified blackbody (MBB) models, characterized by either low spectral indices or low brightness temperatures. Three possible explanations for the deviation are discussed. (1) Optically thick Class 0/I Young stellar objects (YSOs) with very small beam filling factors can lead to lower brightness temperatures than what MBB models predict. (2) Large (mm/cm-sized) dust grains have more significant self-scattering, and therefore frequency-dependent albedo could cause lower spectral indices. (3) Free-free emission over 30 uJy can severely contaminate dust emission and cause low spectral indices for mJy sources in our sample, although the needed number of massive protostars (embedded UCHII regions) is infeasibly high for the normal stellar initial mass function. A reliable measurement of the SFR at low protostellar masses will require future work to distinguish between these possible explanations.

astro-ph.GA

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1)extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2)improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The related code, demos, and pre-trained models are available at https://github.com/jishengpeng/WavTokenizer.

eess.AS

Exploring Text-Queried Sound Event Detection with Audio Source Separation

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22\% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED.

eess.AS

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a significant challenge, especially considering human conversation dynamics such as interruptions, backchannels, and overlapping speech. In this paper, we introduce a novel End-to-End GPT-based model OmniFlatten for full-duplex conversation, capable of effectively modeling the complex behaviors inherent to natural conversations with low latency. To achieve full-duplex conversation capabilities, we propose a multi-stage post-training scheme that progressively adapts a text large language model (LLM) backbone into a speech-text dialogue LLM, capable of generating text and speech in real time, without modifying the architecture of the backbone LLM. The training process comprises three stages: modality alignment, half-duplex dialogue learning, and full-duplex dialogue learning. In all training stages, we standardize the data using a flattening operation, which enables unifying the training methods and the GPT backbone across different modalities and tasks. Our approach offers a simple modeling technique and a promising research direction for developing efficient and natural end-to-end full-duplex spoken dialogue systems. Audio samples of dialogues generated by OmniFlatten can be found at this web site (https://omniflatten.github.io/).

cs.CL

Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision

Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the augmented and original views. Due to lack of negative pairs in the SDPN training process, the network tends to align positive pairs quite closely in the embedding space, a phenomenon known as model collapse. To mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN among self-supervised speaker verification approaches. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1-H, without using any speaker labels in training. Ablation studies show that both proposed learnable prototypes in self-distillation network and diversity regularization contribute to the verification performance.

eess.AS

3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system's proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker.

eess.AS

The deuterium fractionation of NH$_3$ in massive star-forming regions

Deuteration is sensitive to environmental conditions in star-forming regions. To investigate NH$_2$D chemistry, we compared the spatial distribution of ortho-NH$_2$D $1_{11}^s-1_{01}^a$, NH$_3$(1,1) and NH$_3$(2,2) in 12 late-stage massive star-forming regions. By averaging several pixels along the spatial slices of ortho-NH$_2$D $1_{11}^s-1_{01}^a$, we obtained the deuterium fractionation of NH$_3$. In seven targets, the deuterium fractionation of NH$_3$ shows a decreasing trend with increasing rotational temperature. This trend is less clear in the remaining five sources, likely due to limited spatial resolution. However, when considering all 12 sources together, the anticorrelation between NH$_3$ deuterium fractionation and rotational temperature becomes less significant, suggesting that other physical parameters may influence the fractionation. Additionally, we found that the region of highest deuterium fractionation of NH$_3$ is offset from the NH$_3$ peak in each source, likely because the temperature is higher near the NH$_3$ peaks and NH$_2$D may be depleted from the gas phase as the molecular cloud core evolves, as well as the increased release of CO from grains into the gas phase.

astro-ph.GA

OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup

The scaling up has brought tremendous success in the fields of vision and language in recent years. When it comes to audio, however, researchers encounter a major challenge in scaling up the training data, as most natural audio contains diverse interfering signals. To address this limitation, we introduce Omni-modal Sound Separation (OmniSep), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the Query-Mixup strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as Query-Aug, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}.

cs.SD

MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual narratives. This paper presents MuVi, a novel framework that effectively addresses these challenges to enhance the cohesion and immersive experience of audio-visual content. MuVi analyzes video content through a specially designed visual adaptor to extract contextually and temporally relevant features. These features are used to generate music that not only matches the video's mood and theme but also its rhythm and pacing. We also introduce a contrastive music-visual pre-training scheme to ensure synchronization, based on the periodicity nature of music phrases. In addition, we demonstrate that our flow-matching-based music generator has in-context learning ability, allowing us to control the style and genre of the generated music. Experimental results show that MuVi demonstrates superior performance in both audio quality and temporal synchronization. The generated music video samples are available at https://muvi-v2m.github.io.

cs.SD