SearcharxivSearch

arXiv subjects

Kaidi Wang

Publications and source records attributed to Kaidi Wang.

At least 19 recordsLinked to original sources

Pinching-Antenna Systems: From Antenna Placement to Antenna Roaming

This paper investigates pinching-antenna systems with finite antenna movement speed, under which conventional antenna placement is subject to non-negligible repositioning delay, resulting in a fundamental tradeoff between channel quality and effective transmission time. In this context, antenna roaming is proposed as a novel operation mode, in which the antenna moves continuously along the waveguide while simultaneously serving users. By incorporating communication and antenna movement into the same transmission cycle, a unified cycle-duration framework is established to facilitate a fair comparison between antenna placement and antenna roaming. The sum-rate difference is then analytically derived, indicating that antenna roaming avoids dedicated positioning overhead with a loss in channel quality due to spatial averaging. The corresponding sum rate maximization problems are formulated for the two operation modes. The continuous optimization problems are transformed into finite-state sequential decision problems and solved via dynamic programming (DP) based algorithms. For antenna roaming, the optimal unconstrained service-interval partition is analytically characterized, with each antenna position assigned to the user achieving the highest instantaneous rate. Simulation results validate the theoretical analysis, demonstrate the performance advantage of antenna roaming over antenna placement, especially when positioning overhead is significant, and confirm the effectiveness of the DP based solutions in improving the achievable sum rate.

eess.SP

Age-of-Information Aware Federated Learning with Finite Speed Pinching Antenna

This paper investigates age-of-information (AoI) aware federated learning over wireless networks with finite speed pinching antennas. In contrast to existing studies that assume an infinitely high antenna moving speed, a practical round based training procedure is considered, where the pinching antenna is repositioned during the local training phase and its feasible movement range depends on the selected devices. This creates a new coupling among device selection, antenna placement, local training time, model uploading time, and AoI evolution. To characterize the impact of antenna moving speed, the rate gain over the fixed antenna and the gap to the infinite speed benchmark are analyzed. Subsequently, an overall AoI minimization problem is formulated under a round latency deadline by jointly optimizing the selected device set and the pinching antenna position. A coalitional game based device selection algorithm is proposed, where finite speed antenna placement is incorporated into the coalition utility evaluation. For antenna placement, the optimal search region is derived by exploiting the mobility constraint and device location span, based on which a branch-and-bound (BnB) algorithm is developed to obtain the global optimum. Simulation results show that the proposed scheme can accelerate learning convergence, reduce the sum AoI, and improve device participation compared with baseline schemes, demonstrating the potential of pinching antennas for enhancing federated learning through flexible spatial reconfiguration.

eess.SP

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audio but lack linguistic constraints, causing content errors during generation, whereas semantic tokens from self-supervised learning (SSL) models ensure precise text alignment but discard some acoustic information. To bridge this gap, we propose SARA, a dual-stream VAE that directly fuses a frozen SSL semantic anchor with a dedicated residual acoustic encoder. This effectively mitigates the dilemma, creating an efficient and compact latent space without relying on complex regularizers. SARA achieves superior reconstruction quality over strong baselines. Furthermore, in downstream zero-shot TTS tasks, it yields highly natural and expressive synthesis quality, and maintains robust generation performance even under accelerated inference, offering a favorable trade-off between synthesis speed and computational cost.

cs.SD

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

Video dubbing is a cornerstone of multimedia content creation, aiming to synthesize synchronized acoustic sequences for visual streams. While Text-to-Speech (TTS) and Text-to-Audio (TTA) generation have each achieved remarkable progress, existing dubbing systems remain confined to isolated speech synthesis without incorporating sound effects and ambient audio, forcing practitioners to rely on fragmented workflows and laborious manual post-mixing. To address this limitation, we present HoliDubber, a holistic video dubbing framework that moves beyond speech-only generation by enabling the joint synthesis of speech and sound effects from a single text prompt. Specifically, HoliDubber adopts a patch-based autoregressive diffusion transformer architecture, where a causal language model autoregressively models aggregated patch embeddings to capture global temporal structure, and a Diffusion Transformer decoder generates high-fidelity continuous tokens within each patch, following a divide-and-conquer strategy. To achieve cross-modal alignment, visual features are encoded into patch-level representations and fused with audio patches via cross-attention, enabling the model to ground speech generation in the speaker's visual articulation dynamics. In addition, we introduce HoliDub-Bench, a benchmark curated from established datasets with synchronized video-text-audio triplets designed for holistic dubbing evaluation. Extensive experiments demonstrate that HoliDubber significantly outperforms existing methods across multiple benchmarks in speech quality, synchronization, and speaker similarity. Furthermore, results on HoliDub-Bench validate the effectiveness of joint speech-and-sound generation, establishing a new paradigm for holistic video dubbing in complex acoustic scenes. \footnote{The demo page of the project is https://holidubber.github.io}

eess.AS

Energy Efficiency Maximization for Discrete Activation based NOMA-assisted Pinching-Antenna Systems

Pinching-antenna systems is a promising architecture for flexible wireless communications, but energy efficiency (EE) maximization remains largely unexplored, as limited existing studies mainly focus on transmit power minimization. This paper investigates EE maximization in a downlink non-orthogonal multiple access (NOMA)-assisted PASS by explicitly modeling the pinching antenna (PA) activation power and jointly optimizing discrete PA activation and power allocation under both quality-of-service and transmit power constraints. To tackle the resulting mixed-integer nonlinear programming problem, a two-layer iterative algorithm is proposed with an EE-oriented matching-based PA activation and a low-complexity Dinkelbach-based power allocation with closed-form updates. Numerical results demonstrate that the proposed solution achieves substantial EE gains over the considered benchmark schemes, while exhibiting fast convergence. The impact of activation power has been analyzed and the significance of accounting it in EE maximization problem is also demonstrated.

eess.SP

Leaky-Coaxial Pinching-Antenna System with Adjustable Slot Apertures

As a practical physical implementation of pinching-antenna systems, leaky coaxial cable (LCX) enables distributed radiation in more general wireless environments, particularly for lower-frequency applications. In this paper, a leaky-coaxial pinching-antenna system, referred to as the LCX pinching-antenna system, is investigated, and adjustable slot apertures are introduced, such that the slot size can be continuously adjusted rather than being restricted to binary activation. Specifically, the aperture adjustment is modeled as amplitude scaling of the channels induced by the corresponding slots, or equivalently, as power coefficients associated with different slots. Accordingly, analytical results are derived to quantify the performance gain of continuous aperture adjustment over binary slot activation and to reveal the impact of channel coherence on the achievable data rate improvement. Furthermore, static and dynamic time-division multiple access (TDMA) schemes are considered, and the corresponding sum rate maximization problems are formulated and efficiently solved by quadratic transform based optimization, combined with successive convex approximation and alternating updates. Simulation results demonstrate that the proposed design can significantly outperform conventional fixed-antenna systems, traditional LCX schemes, and binary slot activation in terms of both achievable sum rate and outage probability.

eess.SP

Leaky Coaxial Cable based Generalized Pinching-Antenna Systems with Dual-Port Feeding

By leveraging the distributed leakage radiation of leaky coaxial cables (LCXs), the concept of pinching antennas can be generalized from the conventional high-frequency waveguide based architectures to cable based structures in lower-frequency scenarios. This paper investigates an LCX based generalized pinching-antenna system with dual-port feeding. By enabling bidirectional excitation along each cable, the proposed design significantly enhances spatial degrees of freedom. A comprehensive channel model is developed to characterize intra-cable attenuation, bidirectional phase progression, slot based radiation, and wireless propagation. Based on this model, both analog and hybrid beamforming frameworks are studied with the objective of maximizing the minimum achievable data rate. For analog transmission, slot activation, port selection, and power allocation are jointly optimized using matching theory, coalitional games, and bisection based power control. For hybrid transmission, zero-forcing (ZF) digital precoding is incorporated to eliminate inter-user interference, thereby simplifying slot activation and enabling closed-form optimal power allocation. Simulation results demonstrate that dual-port feeding provides notable performance gains over single-port LCX systems and fixed-antenna benchmarks, validating the effectiveness of the proposed beamforming and resource allocation designs under various transmit power levels and cable parameters.

eess.SP

Generalized Pinching-Antenna Systems: A Leaky-Coaxial-Cable Perspective

The evolution toward the sixth-generation (6G) wireless networks has flexible reconfigurable antenna architectures capable of adapting their radiation characteristics to the surrounding environment. At the center stage, while waveguide based pinching antennas have been shown to beneficially ameliorate wireless propagation environments, their applications have remained confined to high-frequency scenarios. As a remedy, we propose a downlink generalized pinching-antenna system that adapts this compelling concept to low-frequency operation through a leaky-coaxial-cable (LCX) implementation. By endowing LCX structures with controllable radiation slots, the system inherits the key capabilities of waveguide based pinching antennas. Explicitly, these include reconfigurable line-of-sight (LoS) links, reduced path loss, and flexible deployment, while supporting a practical implementation of the pinching-antenna concept at low frequencies. A twin-stage propagation model is developed for characterizing both the guided transmission and wireless radiation encountered over LoS and non-line-of-sight (NLoS) paths. Analytical results reveal strong local gain, complemented by rapid distance-dependent decay. Hence, we conceive a matching joint optimization framework, which maximizes throughput by harnessing game theoretic association and convex power allocation. Simulation results demonstrate substantial performance gains over conventional fixed-antenna benchmarks.

eess.SP

SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.

eess.AS

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

Large language models (LLMs) have demonstrated promising performance in both automatic speech recognition (ASR) and text-to-speech (TTS) systems, gradually becoming the mainstream approach. However, most current approaches address these tasks separately rather than through a unified framework. This work aims to integrate these two tasks into one unified model. Although discrete speech tokenization enables joint modeling, its inherent information loss limits performance in both recognition and generation. In this work, we present UniVoice, a unified LLM framework through continuous representations that seamlessly integrates speech recognition and synthesis within a single model. Our approach combines the strengths of autoregressive modeling for speech recognition with flow matching for high-quality generation. To mitigate the inherent divergence between autoregressive and flow-matching models, we further design a dual attention mechanism, which switches between a causal mask for recognition and a bidirectional attention mask for synthesis. Furthermore, the proposed text-prefix-conditioned speech infilling method enables high-fidelity zero-shot voice cloning. Experimental results demonstrate that our method can achieve or exceed current single-task modeling methods in both ASR and zero-shot TTS tasks. This work explores new possibilities for end-to-end speech understanding and generation. Code is available at https://github.com/gwh22/UniVoice.

eess.AS

Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction

Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction.

eess.AS

XMUspeech Systems for the ASVspoof 5 Challenge

In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition.

cs.SD

Pinching-Antenna Systems for Physical Layer Security

This letter investigates the potential of pinching-antenna systems for enhancing physical layer security. By pre-installing multiple pinching antennas at discrete positions along a waveguide, the capability of the considered system to perform amplitude and phase adjustment is validated through the formulation of a secrecy rate maximization problem. Specifically, amplitude control is applied to enhance the signal quality at the legitimate user, while phase alignment is designed to degrade the received signal quality at the eavesdropper. This cooperation among pinching antennas is modeled as a coalitional game, and a corresponding antenna activation algorithm is proposed. The individual impact of each antenna is quantified based on the Shapley value and marginal contribution, providing a fair and efficient method for performance evaluation. Simulation results show that the considered pinching-antenna system achieves significant improvements in secrecy rate, and that the Shapley value based algorithm outperforms conventional coalition value based solutions.

eess.SP

Pinching-Antenna Systems with LoS Blockages

The aim of this letter is to explore the capability of pinching-antenna systems to construct line-of-sight (LoS) links in the presence of LoS blockages. Specifically, pinching antennas are pre-installed at preconfigured positions along waveguides and can be selectively activated to create LoS links for enhancing desired signals and non-line-of-sight (NLoS) links for eliminating inter-user interference. On this basis, a sum-rate maximization problem is formulated by jointly optimizing waveguide assignment and antenna activation. To solve this problem, a matching based algorithm is proposed using two distinct preference designs. Simulation results demonstrate that the considered pinching-antenna system and proposed solutions can dynamically establish LoS links and effectively exploit LoS blockages to mitigate interference, thereby significantly improving system throughput.

eess.SP

A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement

This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins for single-channel speech enhancement. The sub-band module captures surrounding frequency bin information at the input, while the deep filtering module applies filtering at the output to both the target TF bin and its surrounding TF bins. To further improve the model performance, we decouple deep filtering into temporal and frequency components and introduce a two-stage framework, reducing the complexity of filter coefficient prediction at each stage. Additionally, we propose the TAConv module to strengthen convolutional feature extraction. Experimental results demonstrate that the proposed hierarchical deep filtering network (HDF-Net) effectively utilizes surrounding TF bin information and outperforms other advanced systems while using fewer resources.

cs.SD

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization

In recent years, diffusion-based generative models have demonstrated remarkable performance in speech conversion, including Denoising Diffusion Probabilistic Models (DDPM) and others. However, the advantages of these models come at the cost of requiring a large number of sampling steps. This limitation hinders their practical application in real-world scenarios. In this paper, we introduce ReFlow-VC, a novel high-fidelity speech conversion method based on rectified flow. Specifically, ReFlow-VC is an Ordinary Differential Equation (ODE) model that transforms a Gaussian distribution to the true Mel-spectrogram distribution along the most direct path. Furthermore, we propose a modeling approach that optimizes speaker features by utilizing both content and pitch information, allowing speaker features to reflect the properties of the current speech more accurately. Experimental results show that ReFlow-VC performs exceptionally well in small datasets and zero-shot scenarios.

cs.SD

Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.

cs.SD

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech tokenizers has become increasingly important. This paper introduces DS-Codec, a novel neural speech codec featuring a dual-stage training framework with mirror and non-mirror architectures switching, designed to achieve superior speech reconstruction. We conduct extensive experiments and ablation studies to evaluate the effectiveness of our training strategy and compare the performance of the two architectures. Our results show that the mirrored structure significantly enhances the robustness of the learned codebooks, and the training strategy balances the advantages between mirrored and non-mirrored structures, leading to improved high-fidelity speech reconstruction.

cs.SD