SearcharxivSearch

arXiv subjects

Xiang Lv

Publications and source records attributed to Xiang Lv.

At least 19 recordsLinked to original sources

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.

eess.AS

Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-view inconsistency during 3D optimization, as Score Distillation Sampling inherently performs on each single view, inevitably resulting in cross-view hallucinations. To solve above issues, we propose I2C-3D, a novel optimization-based method to generate multi-view consistent compositional 3D assets with reasonable interactions. Specifically, we propose an Inclusive Interactive Collisions strategy to guide Gaussian primitives appearing in reasonable interaction regions naturally, thereby ensuring objects in the compositional scene interact in a physically plausible and visually coherent way. Additionally, to enhance multi-view consistency, Multi-View Adaptive Score Distillation Sampling is devised to distill multi-view consistency prior and layout prior from pre-trained diffusion model by modulating attention map of instance token and spatial token across viewpoints. Benefiting from above elaborate designs, I2C-3D not only generates high-fidelity multi-view consistent compositional 3D assets but also supports 3D editing flexibly, facilitating complex scene generation. Extensive experiments demonstrate our I2C-3D outperforms existing methods in generation quality and multi-view consistency.

cs.CV

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models without auxiliary models. We show that a weighted reward combination converges faster than a probabilistic scheme, and identify three practical optimizations: omitting classifier-free guidance (CFG) during training accelerates convergence; synthesizing hard cases improves robustness; and applying RL to the FM component enhances audio-detail metrics. Experiments on CosyVoice 3.0 and F5-TTS demonstrate objective and subjective preference gains in speaker similarity and perceptual quality, with F5-TTS also improving intelligibility.

eess.AS

Experimental realization of quantum Zeno dynamics for robust quantum metrology

Quantum Zeno dynamics (QZD), which restricts the system's evolution to a protected subspace, provides a promising approach for protecting quantum information from noise. Here, we explore a practical approach to harnessing QZD for robust quantum metrology. By introducing strong inter-particle interactions during the parameter encoding stage, we overcome the typical limitations of previous QZD studies, which have largely focused on single-particle systems and faced challenges where QZD could interfere with the encoding process. We experimentally validate the proposed scheme on a nuclear magnetic resonance platform, achieving near-optimal precision scaling under amplitude damping in both parallel and sequential settings. Numerical simulations further demonstrate the scalability of the approach and its compatibility with other control techniques for suppressing more general types of noise. These findings highlight QZD as a powerful strategy for noise-resilient quantum metrology.

quant-ph

Laboratory formation of scaled astrophysical outflows

Astrophysical systems exhibit a rich diversity of outflow morphologies, yet their mechanisms and existence conditions remain among the most persistent puzzles in the field. Here we present scaled laboratory experiments based on laser-driven plasma outflow into magnetized ambient gas, which mimic five basic astrophysical outflows regulated by interstellar medium, namely collimated jets, blocked jets, elliptical bubbles, as well as spherical winds and bubbles. Their morphologies and existence conditions are found to be uniquely determined by the external Alfvenic and sonic Mach numbers Me-a and Me-s, i.e. the relative strengths of the outflow ram pressure against the magnetic/thermal pressures in the interstellar medium, with transitions occurring at Me-a ~ 2 and 0.5, as well as Me-s ~ 1. These results are confirmed by magnetohydrodynamics simulations and should also be verifiable from existing and future astronomical observations. Our findings provide a quantitative framework for understanding astrophysical outflows.

physics.plasm-ph

Fun-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR .

cs.CL

Generalized Ornstein-Uhlenbeck process for affine stochastic functional differential equations and its applications

This paper studies the existence and global stability of generalized Ornstein-Uhlenbeck process for affine stochastic functional differential equations. Various very basic and important properties are established. In the applications, we present a standard and rigorous procedure for guaranteeing the existence and uniqueness of random equilibria for nonlinear stochastic functional differential equations, which attracts all pull-back trajectories in different types of convergence. Some examples are given to illustrate our main results. The results presented in this paper improve and simplify the conclusions of Jiang and Lv [{\it SIAM J. Control Optim.}, 54 (2016), pp. 2383-2402] and [{\it J. Differential Equations}, 367 (2023), pp. 890-921].

math.DS

An abstract criterion on the existence and global stability of stationary solutions for random dynamical systems and its applications

We prove a concise and easily verifiable criterion on the existence and global stability of stationary solutions for random dynamical systems (RDSs). As a consequence, we can show that the $\omega$-limit sets of all pullback trajectories of semilnear/nonlinear stochastic differential equations (SDEs) with additive/multiplicative white noise are composed of nontrivial random equilibria. The proof is different from the classical RDS scheme, which was established in \cite{CKS}. Furthermore, in the applications of stability analysis for SDEs, our conditions are not only sufficient but indeed sharp.

math.DS

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing E2E approaches primarily fall into two categories: (1) Methods that generate discrete speech tokens independently without incorporating them into the LLM's autoregressive process, resulting in text generation being unaware of concurrent speech synthesis. (2) Models that generate interleaved or parallel speech-text tokens through joint autoregressive modeling, enabling mutual modality awareness during generation. This paper presents DrVoice, a parallel speech-text voice conversation model based on joint autoregressive modeling, featuring dual-resolution speech representations. Notably, while current methods utilize mainly 12.5Hz input audio representation, our proposed dual-resolution mechanism reduces the input frequency for the LLM to 5Hz, significantly reducing computational cost and alleviating the frequency discrepancy between speech and text tokens and in turn better exploiting LLMs' capabilities. Experimental results demonstrate that DrVoice-7B establishes new state-of-the-art (SOTA) on prominent speech benchmarks including OpenAudioBench, VoiceBench, UltraEval-Audio and Big Bench Audio, making it a leading open-source speech foundation model in ~7B models.

cs.CL

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

Modern robot navigation systems encounter difficulties in diverse and complex indoor environments. Traditional approaches rely on multiple modules with small models or rule-based systems and thus lack adaptability to new environments. To address this, we developed Astra, a comprehensive dual-model architecture, Astra-Global and Astra-Local, for mobile robot navigation. Astra-Global, a multimodal LLM, processes vision and language inputs to perform self and goal localization using a hybrid topological-semantic graph as the global map, and outperforms traditional visual place recognition methods. Astra-Local, a multitask network, handles local path planning and odometry estimation. Its 4D spatial-temporal encoder, trained through self-supervised learning, generates robust 4D features for downstream tasks. The planning head utilizes flow matching and a novel masked ESDF loss to minimize collision risks for generating local trajectories, and the odometry head integrates multi-sensor inputs via a transformer encoder to predict the relative pose of the robot. Deployed on real in-house mobile robots, Astra achieves high end-to-end mission success rate across diverse indoor environments.

cs.RO

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.

cs.SD

Fisher-Based Sensitivity Framework for Rydberg Atom Microwave Electrometry

Fisher information provides a rigorous theoretical benchmark for evaluating quantum sensor sensitivity; however, a comprehensive framework for quantifying the fundamental limits of Rydberg-atom microwave electrometers remains lacking. In this work, we establish such a framework by deriving the Fisher information for slope detection and establishing its connection to sensitivity through signal-to-noise ratio, leading to an analytical expression jointly determined by photon shot noise and atomic response. Numerical implementation with real parameters in cesium vapor systems reveals a Fisher-optimized sensitivity below $\mathrm{nV\,cm^{-1}\,Hz^{-1/2}}$, highlighting a substantial potential for sensitivity enhancement in practical experiments through the suppression of technical noise. Importantly, the theory predicts that sub-nanovolt sensitivity is robust against moderate variations in system parameters, thereby delineating both the ultimate sensitivity and optimal operational regime of Rydberg-atom microwave electrometers.

physics.optics

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice interactions are categorized as native and aligned. Native models integrate speech and text processing in one framework but struggle with issues like differing sequence lengths and insufficient pre-training. Aligned models maintain text LLM capabilities but are often limited by small datasets and a narrow focus on speech tasks. In this work, we introduce MinMo, a Multimodal Large Language Model with approximately 8B parameters for seamless voice interaction. We address the main limitations of prior aligned multimodal models. We train MinMo through multiple stages of speech-to-text alignment, text-to-speech alignment, speech-to-speech alignment, and duplex interaction alignment, on 1.4 million hours of diverse speech data and a broad range of speech tasks. After the multi-stage training, MinMo achieves state-of-the-art performance across various benchmarks for voice comprehension and generation while maintaining the capabilities of text LLMs, and also facilitates full-duplex conversation, that is, simultaneous two-way communication between the user and the system. Moreover, we propose a novel and simple voice decoder that outperforms prior models in voice generation. The enhanced instruction-following capabilities of MinMo supports controlling speech generation based on user instructions, with various nuances including emotions, dialects, and speaking rates, and mimicking specific voices. For MinMo, the speech-to-text latency is approximately 100ms, full-duplex latency is approximately 600ms in theory and 800ms in practice. The MinMo project web page is https://funaudiollm.github.io/minmo, and the code and models will be released soon.

cs.CL

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

In our previous work, we introduced CosyVoice, a multilingual speech synthesis model based on supervised discrete speech tokens. By employing progressive semantic decoding with two popular generative models, language models (LMs) and Flow Matching, CosyVoice demonstrated high prosody naturalness, content consistency, and speaker similarity in speech in-context learning. Recently, significant progress has been made in multi-modal large language models (LLMs), where the response latency and real-time factor of speech synthesis play a crucial role in the interactive experience. Therefore, in this report, we present an improved streaming speech synthesis model, CosyVoice 2, which incorporates comprehensive and systematic optimizations. Specifically, we introduce finite-scalar quantization to improve the codebook utilization of speech tokens. For the text-speech LM, we streamline the model architecture to allow direct use of a pre-trained LLM as the backbone. In addition, we develop a chunk-aware causal flow matching model to support various synthesis scenarios, enabling both streaming and non-streaming synthesis within a single model. By training on a large-scale multilingual dataset, CosyVoice 2 achieves human-parity naturalness, minimal response latency, and virtually lossless synthesis quality in the streaming mode. We invite readers to listen to the demos at https://funaudiollm.github.io/cosyvoice2.

cs.SD

Transverse Bending Mimicry of Longitudinal Piezoelectricity

The origin of frequently observed ultrahigh electric-induced longitudinal strain, ranging from 1% to 26%, remains an open question. Recent evidence suggests that this phenomenon is linked to the bending deformation of samples, but the mechanisms driving this bending and the strong dependence of nominal strain on sample thickness have yet to be fully understood. Here, we demonstrate that the bending in piezoceramics can be induced by non-zero gradient of d31 acrcoss thickness direction. Our calculations show that in standard perovskite piezoceramics, such as KNbO3, a 0.69% concentration of oxygen vacancies results in a 6.3 pC/N change in d31 by inhibiting polarization rotation, which is sufficient to produce ultrahigh nominal strain in thin samples. The gradients of defect concentration, composition, and stress can all cause sufficient inhomogeneity in the distribution of d31, leading to the bending effect. We propose several approaches to distinguish true electric-induced strain from bending-induced effects. Our work provides clarity on the origin of nominal ultrahigh electricinduced strain and offers valuable insights for advancing piezoelectric materials.

cond-mat.mtrl-sci

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https://fun-audio-llm.github.io, and the code can be accessed at https://github.com/FunAudioLLM.

cs.SD

The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge

We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-Channel Multi-Party Meeting Transcription (M2MeT) challenge. These techniques are designed to handle multi-speaker conversations in real-world meeting scenarios with high speaker-overlap ratios and under heavy reverberant and noisy condition. First, for data preparation and augmentation in training TS-VAD models, speech data containing both real meetings and simulated indoor conversations are used. Second, in refining results obtained after TS-VAD based decoding, we perform a series of post-processing steps to improve the VAD results needed to reduce diarization error rates (DERs). Tested on the ALIMEETING corpus, the newly released Mandarin meeting dataset used in M2MeT, we demonstrate that our proposed system can decrease the DER by up to 66.55/60.59% relatively when compared with classical clustering based diarization on the Eval/Test set.

eess.AS

Possible implications for particle physics by quantum measurement

In sharp contrast to its classical counterpart, quantum measurement plays a fundamental role in quantum mechanics and blurs the essential distinction between the measurement apparatus and the objects under investigation. An appealing phenomenon in quantum measurements, termed as quantum Zeno effect, can be observed in particular subspaces selected by measurement Hamiltonian. Here we apply the top-down Zeno mechanism to the particle physics. We indeed develop an alternative insight into the properties of fundamental particles, but not intend to challenge the Standard Model (SM). In a unified and simple manner, our effective model allows to merge the origin of neutrino's small mass and oscillations, the hierarchy pattern for masses of electric charged fermions, the color confinement, and the discretization of quantum numbers, using a perturbative theory for the dynamical quantum Zeno effect. Under various conditions for vanishing transition amplitudes among particle eigenstates in the effective model, it is remarkable to probe results that are somewhat reminiscent of SM, including: (i) neutrino oscillations with big-angle mixing and small masses emerge from the energy-momentum conservation, (ii) electrically-charged fermions hold masses in a hierarchy pattern due to the electric-charge conservation, (iii) color confinement and the associated asymptotic freedom can be deduced from the color-charge conservation. We make several anticipations about the basic properties for fundamental particles: (i) the total mass of neutrinos and the existence of a nearly massless neutrino (of any generation), (ii) the discretization in quantum numbers for the new-discovered electrically-charged fermions, (iii) the confinement and the associated asymptotic freedom for any particle containing more than two conserved charges.

hep-ph