SearcharxivSearch

arXiv subjects

Qiaoling Huang

Publications and source records attributed to Qiaoling Huang.

2 recordsLinked to original sources

StepAudio 2.5 Technical Report

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to match the depth of specialized systems across automatic speech recognition (ASR), text-to-speech synthesis (TTS), and realtime spoken interaction. Bridging this gap remains an open challenge. This report presents StepAudio 2.5, a unified audio-language foundation model that matches or exceeds specialized systems across all three capabilities. Rather than treating these tasks as architecturally distinct, we operate on the premise that once text and audio share a multimodal representational space, task specialization becomes a matter of operational regimes: data construction, optimization targets, and decoding constraints. Guided by this insight, we advance the post-training paradigm from standard supervised learning to task-tailored Reinforcement Learning from Human Feedback (RLHF), using it as the primary mechanism to define complex optimization targets. We leverage this RLHF-centric alignment, alongside specialized decoding, to shape a shared backbone into three distinct operational modes. Concretely, the ASR branch advances transcription efficiency via verifiable multi-token decoding; the TTS branch achieves controllable, expressive synthesis through preference-based RLHF and context-rich supervision; and the Realtime branch realizes low-latency, persona-consistent dialogue via generative reward modeling within an RLHF framework. On standard benchmarks, StepAudio 2.5 achieves state-of-the-art results across ASR, TTS, and Realtime, demonstrating that a singular audio-language foundation can successfully internalize the distinct deployment objectives of speech understanding, generation, and live interaction.

eess.AS

Raman spectroscopy of optical phonon and incommensurate charge density wave modes in 2H-TaSe2 exfoliated flakes

2H-TaSe2 is a model transition metal dichalcogenide material that develops charge density waves (CDWs).Here we present variable-temperature Raman spectroscopy study on both incommensurate charge density waves (ICDW) and optical phonon modes of 2H-TaSe2 thin layers exfoliated onto SiO2 substrate. Raman scattering intensities of all modes reach a maximum when the sample thickness is about 11 nm. This phenomenon can be explained by optical interference effect between the sample and the substrate. The E2gICDW amplitude modes experience redshift as temperature increases. We extract ICDW transition temperature (TICDW) from temperature dependence of the frequency of E2gICDW mode. We find that TICDW increases in thinner flakes,which could be due to a result of significantly enhanced electron-phonon interactions. Our results open up a new window for search and control of CDW of two-dimensional matter.

cond-mat.mtrl-sci