SearcharxivSearch

arXiv subjects

Wenchao Wang

Publications and source records attributed to Wenchao Wang.

At least 19 recordsLinked to original sources

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.

cs.SD

Highly-efficient perturbative Raman shifting by engineering the nonlinear temporal response

Raman scattering underlies a broad range of spectroscopic and light-generation techniques, yet its conventional description, based on the Raman gain spectrum, accurately describes only long-pulse, steady-state dynamics. We present a time-domain theoretical approach that provides a unified and physically-transparent description of Raman interactions across all temporal regimes. It enables direct visualization of Raman temporal dynamics and accounts for spectrotemporal aspects of Raman phenomena, which cannot be addressed by prior theories. In particular, molecules with strong Raman responses do not produce an efficient soliton self-frequency shift in gas-filled hollow-core fibers. The time-domain analysis exposes temporal and spectral distortions from the Raman response that impact frequency-shifting detrimentally, and identifies how these distortions can be suppressed by reducing the Raman interaction to a perturbation on the electronic response. Experiments that employ gas mixtures with tunable Raman fractions of the nonlinear response demonstrate up to a four-fold increase in quantum efficiency (from 20 to 80%) compared to the pure molecular gas, and unity-efficiency Raman shifting will be possible. The new time-domain framework uncovers phenomena that are inaccessible through the decades-old frequency-domain treatment of Raman scattering, and it applies to Raman interactions in solids, liquids, and gases on various timescales.

physics.optics

Partial parabolic amplification in rare-earth-doped optical fiber

Nonlinear amplification is a powerful technique for generating ultrashort laser pulses with high peak power in fiber systems. However, the diversity of nonlinear amplification approaches and their inherent complexities present significant challenges to achieving a unified understanding and further scaling of peak power and pulse energy while preserving ultrashort durations. Here, we report the results of a systematic optimization with respect to seed pulse duration that elucidates the dynamics of nonlinear amplification and allows identification of distinct propagation regimes. As part of this analysis, we identify a new regime, termed partial parabolic amplification, which achieves 50-fs pulse duration and yields higher peak power than any other nonlinear amplification regime known to date. An initial experimental demonstration of partial parabolic amplification produces 50-fs and 2.2-uJ pulses with a 25-um-core Yb fiber amplifier, corresponding to a 30-MW peak power. In contrast to other nonlinear amplification techniques, practical energy scaling beyond 10 uJ and 200 MW should be achievable with available gain fibers with larger mode areas, which would fill a gap in existing fiber laser capabilities that would directly impact material processing, nonlinear bio-imaging, and other applications.

physics.optics

JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice

cs.SD

Confined tumbling state as the origin of the excess of slowly rotating asteroids

The rotational distribution of asteroids as a function of their size is used {as a diagnostic of} their physical properties and evolution. Recent photometric surveys from the Gaia mission, allowing observation of asteroids with long spin periods (for example $\geq 24$h), found an excessive group of slow rotators and a gap separating them from faster rotators, which is unexplained by current theories. Here we developed an asteroid rotational evolution model capable of reproducing the observed distribution. {We suggest that this distribution is regulated by the competition between collisions and internal friction dampening of "tumblers" -asteroids with unstable rotation vectors, and that the slow rotator group is mainly populated by tumblers.} {We constrain the product of the rigidity and quality factor, which relates to the body's viscosity, to $μQ \sim 4 \times 10^9~$Pa. This number, two orders of magnitude smaller than the one assumed for monolithic boulders,} implies that {rubble pile} asteroids could have a porous structure or a thick regolith layer, and undergo stronger tidal effects.

astro-ph.EP

Efficient, broadly-tunable source of megawatt pulses for multiphoton microscopy based on self-phase modulation in argon-filled hollow-core fiber

An exciting recent development for deep-tissue imaging with cellular resolution is three-photon fluorescence microscopy (3PM) with excitation at long wavelengths (1300 and 1700 nm). In the last few years, long-wavelength 3PM has driven rapid progress in deep-tissue imaging beyond the depth limit of two-photon microscopy, with impacts in neuroscience, immunology, and cancer biology. However, wide adoption of 3PM faces challenges. Three-photon excitation (3PE) is naturally weaker than two-photon excitation, which places a premium on ultrashort pulses with high peak power. The inefficiency, complexity, and cost of current sources of these pulses present major barriers to the use of 3PM in typical biomedical research labs. Here, we describe a fiber-based source of femtosecond pulses with multi-megawatt peak power, tunable from 850 nm to 1700 nm. Compressed pulses from a fiber amplifier at 1030~nm are launched into an antiresonant hollow-core fiber filled with argon. By varying only the gas pressure, pulses with hundreds of nanojoules of energy and sub-100 fs duration are obtained at wavelengths between 850 and 1700 nm. This approach is a new route to an efficient, robust, and potentially low-cost source for multiphoton deep-tissue imaging. In particular, 960-nJ and 50-fs pulses are generated at 1300 nm with a conversion efficiency of 10\%. The nearly 20-MW peak power is an order of magnitude higher than the previous best from femtosecond fiber source at 1300~nm. As an example of the capabilities of the source, these pulses are used to image structure and neuronal activity in mouse brain as deep as 1.1 mm below the dura.

physics.optics

Underwater Acoustic Target Recognition based on Smoothness-inducing Regularization and Spectrogram-based Data Augmentation

Underwater acoustic target recognition is a challenging task owing to the intricate underwater environments and limited data availability. Insufficient data can hinder the ability of recognition systems to support complex modeling, thus impeding their advancement. To improve the generalization capacity of recognition models, techniques such as data augmentation have been employed to simulate underwater signals and diversify data distribution. However, the complexity of underwater environments can cause the simulated signals to deviate from real scenarios, resulting in biased models that are misguided by non-true data. In this study, we propose two strategies to enhance the generalization ability of models in the case of limited data while avoiding the risk of performance degradation. First, as an alternative to traditional data augmentation, we utilize smoothness-inducing regularization, which only incorporates simulated signals in the regularization term. Additionally, we propose a specialized spectrogram-based data augmentation strategy, namely local masking and replicating (LMR), to capture inter-class relationships. Our experiments and visualization analysis demonstrate the superiority of our proposed strategies.

cs.SD

Improving Short Utterance Anti-Spoofing with AASIST2

The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during short utterance evaluation. To solve this problem, AASIST can be improved to AASIST2 by modifying the residual blocks to Res2Net blocks. The modified Res2Net blocks can extract multi-scale features and improve the detection performance for speech of different durations, thus improving the short utterance evaluation performance. On the other hand, adaptive large margin fine-tuning (ALMFT) has achieved performance improvement in short utterance speaker verification. Therefore, we apply Dynamic Chunk Size (DCS) and ALMFT training strategies in speech anti-spoofing to further improve the performance of short utterance evaluation. Experiments demonstrate that the proposed AASIST2 improves the performance of short utterance evaluation while maintaining the performance of regular evaluation on different datasets.

eess.AS

Enhancing Spoofing Speech Detection Using Rhythm Information

Current spoofing speech detection systems need more convincing evidence. In this paper, the flaws of rhythm information inherent in the TTS-generated speech are analyzed to increase the reliability of detection systems. TTS models take text as input and utilize acoustic models to predict rhythm information, which introduces artifacts in the rhythm information. By filtering out vocal tract response, the remaining glottal flow with rhythm information retains detection ability for TTS-generated speech. Based on these analyses, a rhythm perturbation module is proposed to enhance the copy-synthesis data augmentation method. Fake utterances generated by the proposed method force the detecting model to pay attention to the artifacts in rhythm information and effectively improve the ability to detect TTS-generated speech of the anti-spoofing countermeasures.

eess.AS

Synthetic Speech Detection Based on Temporal Consistency and Distribution of Speaker Features

Current synthetic speech detection (SSD) methods perform well on certain datasets but still face issues of robustness and interpretability. A possible reason is that these methods do not analyze the deficiencies of synthetic speech. In this paper, the flaws of the speaker features inherent in the text-to-speech (TTS) process are analyzed. Differences in the temporal consistency of intra-utterance speaker features arise due to the lack of fine-grained control over speaker features in TTS. Since the speaker representations in TTS are based on speaker embeddings extracted by encoders, the distribution of inter-utterance speaker features differs between synthetic and bonafide speech. Based on these analyzes, an SSD method based on temporal consistency and distribution of speaker features is proposed. On one hand, modeling the temporal consistency of intra-utterance speaker features can aid speech anti-spoofing. On the other hand, distribution differences in inter-utterance speaker features can be utilized for SSD. The proposed method offers low computational complexity and performs well in both cross-dataset and silence trimming scenarios.

eess.AS

The Impact of Silence on Speech Anti-Spoofing

The current speech anti-spoofing countermeasures (CMs) show excellent performance on specific datasets. However, removing the silence of test speech through Voice Activity Detection (VAD) can severely degrade performance. In this paper, the impact of silence on speech anti-spoofing is analyzed. First, the reasons for the impact are explored, including the proportion of silence duration and the content of silence. The proportion of silence duration in spoof speech generated by text-to-speech (TTS) algorithms is lower than that in bonafide speech. And the content of silence generated by different waveform generators varies compared to bonafide speech. Then the impact of silence on model prediction is explored. Even after retraining, the spoof speech generated by neural network based end-to-end TTS algorithms suffers a significant rise in error rates when the silence is removed. To demonstrate the reasons for the impact of silence on CMs, the attention distribution of a CM is visualized through class activation mapping (CAM). Furthermore, the implementation and analysis of the experiments masking silence or non-silence demonstrates the significance of the proportion of silence duration for detecting TTS and the importance of silence content for detecting voice conversion (VC). Based on the experimental results, improving the robustness of CMs against unknown spoofing attacks by masking silence is also proposed. Finally, the attacks on anti-spoofing CMs through concatenating silence, and the mitigation of VAD and silence attack through low-pass filtering are introduced.

eess.AS

One-Class Knowledge Distillation for Spoofing Speech Detection

The detection of spoofing speech generated by unseen algorithms remains an unresolved challenge. One reason for the lack of generalization ability is traditional detecting systems follow the binary classification paradigm, which inherently assumes the possession of prior knowledge of spoofing speech. One-class methods attempt to learn the distribution of bonafide speech and are inherently suited to the task where spoofing speech exhibits significant differences. However, training a one-class system using only bonafide speech is challenging. In this paper, we introduce a teacher-student framework to provide guidance for the training of a one-class model. The proposed one-class knowledge distillation method outperforms other state-of-the-art methods on the ASVspoof 21DF dataset and InTheWild dataset, which demonstrates its superior generalization ability.

eess.AS

Progressive Sub-Graph Clustering Algorithm for Semi-Supervised Domain Adaptation Speaker Verification

Utilizing the large-scale unlabeled data from the target domain via pseudo-label clustering algorithms is an important approach for addressing domain adaptation problems in speaker verification tasks. In this paper, we propose a novel progressive subgraph clustering algorithm based on multi-model voting and double-Gaussian based assessment (PGMVG clustering). To fully exploit the relationships among utterances and the complementarity among multiple models, our method constructs multiple k-nearest neighbors graphs based on diverse models and generates high-confidence edges using a voting mechanism. Further, to maximize the intra-class diversity, the connected subgraph is utilized to obtain the initial pseudo-labels. Finally, to prevent disastrous clustering results, we adopt an iterative approach that progressively increases k and employs a double-Gaussian based assessment algorithm to decide whether merging sub-classes.

cs.SD

PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker Verification

ECAPA-TDNN is currently the most popular TDNN-series model for speaker verification, which refreshed the state-of-the-art(SOTA) performance of TDNN models. However, one-dimensional convolution has a global receptive field over the feature channel. It destroys the time-frequency relevance of the spectrogram. Besides, as ECAPA-TDNN only has five layers, a much shallower structure compared to ResNet restricts the capability to generate deep representations. To further improve ECAPA-TDNN, we propose a progressive channel fusion strategy that splits the spectrogram across the feature channel and gradually expands the receptive field through the network. Secondly, we enlarge the model by extending the depth and adding branches. Our proposed model achieves EER with 0.718 and minDCF(0.01) with 0.0858 on vox1o, relatively improved 16.1\% and 19.5\% compared with ECAPA-TDNN-large.

cs.SD

Deepfake Detection System for the ADD Challenge Track 3.2 Based on Score Fusion

This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light convolutional neural network (LCNN) based models. Various front-ends are used as input features, including low-frequency short-time Fourier transform and Constant Q transform. Due to the complex noise and rich synthesis algorithms, it is difficult to obtain the desired performance using the training set directly. Online data augmentation methods effectively improve the robustness of fake audio detection systems. In particular, the reasons for the poor improvement of score fusion are explored through visualization of the score distributions and comparison with score distribution on another dataset. The overfitting of the model to the training set leads to extreme values of the scores and low correlation of the score distributions, which makes score fusion difficult. Fusion with partially fake audio detection system improves system performance further. The submission on track 3.2 obtained the weighted equal error rate (WEER) of 11.04\%, which is one of the best performing systems in the challenge.

eess.AS

The HCCL System for the NIST SRE21

This paper describes the systems developed by the HCCL team for the NIST 2021 speaker recognition evaluation (NIST SRE21).We first explore various state-of-the-art speaker embedding extractors combined with a novel circle loss to obtain discriminative deep speaker embeddings. Considering that cross-channel and cross-linguistic speaker recognition are the key challenges of SRE21, we introduce several techniques to reduce the cross-domain mismatch. Specifically, Codec and speech enhancement are directly applied to the raw speech to eliminate the codecs and the environment noise mismatch. We denote the methods that work directly on speech to eliminate the relatively explicit mismatches collectively as data adaptation methods. Experiments show that data adaption methods achieve 15\% improvements over our baseline. Furthermore, some popular back-ends domain adaptation algorithms are deployed on speaker embeddings to alleviate speaker performance degradation caused by the implicit mismatch. Score calibration is a major failure for us in SRE21. The reason is that score calibration with too many parameters easily lead to overfitting problems.

cs.SD

SASV Based on Pre-trained ASV System and Integrated Scoring Module

Based on the assumption that there is a correlation between anti-spoofing and speaker verification, a Total-Divide-Total integrated Spoofing-Aware Speaker Verification (SASV) system based on pre-trained automatic speaker verification (ASV) system and integrated scoring module is proposed and submitted to the SASV 2022 Challenge. The training and scoring of ASV and anti-spoofing countermeasure (CM) in current SASV systems are relatively independent, ignoring the correlation. In this paper, by leveraging the correlation between the two tasks, an integrated SASV system can be obtained by simply training a few more layers on the basis of the baseline pre-trained ASV subsystem. The features in pre-trained ASV system are utilized for logical access spoofing speech detection. Further, speaker embeddings extracted by the pre-trained ASV system are used to improve the performance of the CM. The integrated scoring module takes the embeddings of the ASV and anti-spoofing branches as input and preserves the correlation between the two tasks through matrix operations to produce integrated SASV scores. Submitted primary system achieved equal error rate (EER) of 3.07\% on the development dataset of the SASV 2022 Challenge and 4.30\% on the evaluation part, which is a 25\% improvement over the baseline systems.

eess.AS