SearcharxivSearch

arXiv subjects

Chunshan Liu

Publications and source records attributed to Chunshan Liu.

11 recordsLinked to original sources

InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches based on visual features face two key limitations: (1)they rely on complex, handcrafted visual preprocessing pipelines, including facial landmark detection and feature extraction; and (2) they generalize poorly to unseen visual domains, often resulting in degraded alignment and dubbing quality. To address these issues, we propose InstructDubber, a novel instruction-based alignment dubbing method for both robust in-domain and zero-shot movie dubbing. Specifically, we first feed the video, script, and corresponding prompts into a multimodal large language model to generate natural language dubbing instructions regarding the speaking rate and emotion state depicted in the video, which is robust to visual domain variations. Second, we design an instructed duration distilling module to mine discriminative duration cues from speaking rate instructions to predict lip-aligned phoneme-level pronunciation duration. Third, for emotion-prosody alignment, we devise an instructed emotion calibrating module, which finetunes an LLM-based instruction analyzer using ground truth dubbing emotion as supervision and predicts prosody based on the calibrated emotion analysis. Finally, the predicted duration and prosody, together with the script, are fed into the audio decoder to generate video-aligned dubbing. Extensive experiments on three major benchmarks demonstrate that InstructDubber outperforms state-of-the-art approaches across both in-domain and zero-shot scenarios.

cs.SD

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated prosody structures to build a high-quality video-synchronized dubbing track. The limited scale of movie dubbing datasets, along with the background noise inherent in audio data, hinder the acoustic modeling performance of trained models. To address these issues, we propose an acoustic-prosody disentangled two-stage method to achieve high-quality dubbing generation with precise prosody alignment. First, we propose a prosody-enhanced acoustic pre-training to develop robust acoustic modeling capabilities. Then, we freeze the pre-trained acoustic system and design a disentangled framework to model prosodic text features and dubbing style while maintaining acoustic quality. Additionally, we incorporate an in-domain emotion analysis module to reduce the impact of visual domain shifts across different movies, thereby enhancing emotion-prosody alignment. Extensive experiments show that our method performs favorably against the state-of-the-art models on two primary benchmarks. The demos are available at https://zzdoog.github.io/ProDubber/.

cs.SD

Bayesian Functional Graphical Models with Change-Point Detection

Functional data analysis, which models data as realizations of random functions over a continuum, has emerged as a useful tool for time series data. Often, the goal is to infer the dynamic connections (or time-varying conditional dependencies) among multiple functions or time series. For this task, a dynamic and Bayesian functional graphical model is introduced. The proposed modeling approach prioritizes the careful definition of an appropriate graph to identify both time-invariant and time-varying connectivity patterns. A novel block-structured sparsity prior is paired with a finite basis expansion, which together yield effective shrinkage and graph selection with efficient computations via a Gibbs sampling algorithm. Crucially, the model includes (one or more) graph changepoints, which are learned jointly with all model parameters and incorporate graph dynamics. Simulation studies demonstrate excellent graph selection capabilities, with significant improvements over competing methods. The proposed approach is applied to study of dynamic connectivity patterns of sea surface temperatures in the Pacific Ocean and reveals meaningful edges.

stat.ME

Joint Beam Training and Data Transmission Design for Covert Millimeter-Wave Communication

Covert communication prevents legitimate transmission from being detected by a warden while maintaining certain covert rate at the intended user. Prior works have considered the design of covert communication over conventional low-frequency bands, but few works so far have explored the higher-frequency millimeter-wave (mmWave) spectrum. The directional nature of mmWave communication makes it attractive for covert transmission. However, how to establish such directional link in a covert manner in the first place remains as a significant challenge. In this paper, we consider a covert mmWave communication system, where legitimate parties Alice and Bob adopt beam training approach for directional link establishment. Accounting for the training overhead, we develop a new design framework that jointly optimizes beam training duration, training power and data transmission power to maximize the effective throughput of Alice-Bob link while ensuring the covertness constraint at warden Willie is met. We further propose a dual-decomposition successive convex approximation algorithm to solve the problem efficiently. Numerical studies demonstrate interesting tradeoff among the key design parameters considered and also the necessity of joint design of beam training and data transmission for covert mmWave communication.

cs.IT

Robust Adaptive Beam Tracking for Mobile Millimetre Wave Communications

Millimetre wave (mmWave) beam tracking is a challenging task because tracking algorithms are required to provide consistent high accuracy with low probability of loss of track and minimal overhead. To meet these requirements, we propose in this paper a new analog beam tracking framework namely Adaptive Tracking with Stochastic Control (ATSC). Under this framework, beam direction updates are made using a novel mechanism based on measurements taken from only two beam directions perturbed from the current data beam. To achieve high tracking accuracy and reliability, we provide a systematic approach to jointly optimise the algorithm parameters. The complete framework includes a method for adapting the tracking rate together with a criterion for realignment (perceived loss of track). ATSC adapts the amount of tracking overhead that matches well to the mobility level, without incurring frequent loss of track, as verified by an extensive set of experiments under both representative statistical channel models as well as realistic urban scenarios simulated by ray-tracing software. In particular, numerical results show that ATSC can track dominant channel directions with high accuracy for vehicles moving at 72 km/hour in complicated urban scenarios, with an overhead of less than 1\%.

eess.SP

Millimeter-Wave Beam Search with Iterative Deactivation and Beam Shifting

Millimeter Wave (mmWave) communications rely on highly directional beams to combat severe propagation loss. In this paper, an adaptive beam search algorithm based on spatial scanning, called Iterative Deactivation and Beam Shifting (IDBS), is proposed for mmWave beam alignment. IDBS does not require advance information such as the Signal-to-Noise Ratio (SNR) and channel statistics, and matches the training overhead to the unknown SNR to achieve satisfactory performance. The algorithm works by gradually deactivating beams using a Bayesian probability criterion based on a uniform improper prior, where beam deactivation can be implemented with low-complexity operations that require computing a low-degree polynomial or a search through a look-up table. Numerical results confirm that IDBS adapts to different propagation scenarios such as line-of-sight and non-line-of-sight and to different SNRs. It can achieve better tradeoffs between training overhead and beam alignment accuracy than existing non-adaptive algorithms that have fixed training overheads.

cs.IT

Explore and Eliminate: Optimized Two-Stage Search for Millimeter-Wave Beam Alignment

Swift and accurate alignment of transmitter (Tx) and receiver (Rx) beams is a fundamental design challenge to enable reliable outdoor millimeter-wave communications. In this paper, we propose a new Optimized Two-Stage Search (OTSS) algorithm for Tx-Rx beam alignment via spatial scanning. In contrast to one-shot exhaustive search, OTSS judiciously divides the training energy budget into two stages. In the first stage, OTSS explores and trains all candidate beam pairs and then eliminates a set of less favorable pairs learned from the received signal profile. In the second stage, OTSS takes an extra measurement for each of the survived pairs and combines with the previous measurement to determine the best one. For OTSS, we derive an upper bound on its misalignment probability, under a single-path channel model with training codebooks having an ideal beam pattern. We also characterize the decay rate function of the upper bound with respect to the training budget and further derive the optimal design parameters of OTSS that maximize the decay rate. OTSS is proved to asymptotically outperform state-of-the-art beam alignment algorithms, and is numerically shown to achieve better performance with limited training budget and practically synthesized beams.

cs.IT

Millimeter Wave Beam Alignment: Large Deviations Analysis and Design Insights

In millimeter wave cellular communication, fast and reliable beam alignment via beam training is crucial to harvest sufficient beamforming gain for the subsequent data transmission. In this paper, we establish fundamental limits in beam-alignment performance under both the exhaustive search and the hierarchical search that adopts multi-resolution beamforming codebooks, accounting for time-domain training overhead. Specifically, we derive lower and upper bounds on the probability of misalignment for an arbitrary level in the hierarchical search, based on a single-path channel model. Using the method of large deviations, we characterize the decay rate functions of both bounds and show that the bounds coincide as the training sequence length goes large. We go on to characterize the asymptotic misalignment probability of both the hierarchical and exhaustive search, and show that the latter asymptotically outperforms the former, subject to the same training overhead and codebook resolution. We show via numerical results that this relative performance behavior holds in the non-asymptotic regime. Moreover, the exhaustive search is shown to achieve significantly higher worst-case spectrum efficiency than the hierarchical search, when the pre-beamforming signal-to-noise ratio (SNR) is relatively low. This study hence implies that the exhaustive search is more effective for users situated further from base stations, as they tend to have low SNR.

cs.IT

Precoding for the Sparsely Spread MC-CDMA Downlink with Discrete-Alphabet Inputs

Sparse signatures have been proposed for the CDMA uplink to reduce multi-user detection complexity, but they have not yet been fully exploited for its downlink counterpart. In this work, we propose a Multi-Carrier CDMA (MC-CDMA) downlink communication, where regular sparse signatures are deployed in the frequency domain. Taking the symbol detection point of view, we formulate a problem appropriate for the downlink with discrete alphabets as inputs. The solution to the problem provides a power-efficient precoding algorithm for the base station, subject to minimum symbol error probability (SEP) requirements at the mobile stations. In the algorithm, signature sparsity is shown to be crucial for reducing precoding complexity. Numerical results confirm system-load-dependent power reduction gain from the proposed precoding over the zero-forcing precoding and the regularized zero-forcing precoding with optimized regularization parameter under the same SEP targets. For a fixed system load, it is also demonstrated that sparse MC-CDMA with a proper choice of sparsity level attains almost the same power efficiency and link throughput as that of dense MC-CDMA yet with reduced precoding complexity, thanks to the sparse signatures.

cs.IT

Design and Analysis of Transmit Beamforming for Millimetre Wave Base Station Discovery

In this paper, we develop an analytical framework for the initial access (a.k.a. Base Station (BS) discovery) in a millimeter-wave (mm-wave) communication system and propose an effective strategy for transmitting the Reference Signals (RSs) used for BS discovery. Specifically, by formulating the problem of BS discovery at User Equipments (UEs) as hypothesis tests, we derive a detector based on the Generalised Likelihood Ratio Test (GLRT) and characterise the statistical behaviour of the detector. The theoretical results obtained allow analysis of the impact of key system parameters on the performance of BS discovery, and show that RS transmission with narrow beams may not be helpful in improving the overall BS discovery performance due to the cost of spatial scanning. Using the method of large deviations, we identify the desirable beam pattern that minimises the average miss-discovery probability of UEs within a targeted detectable region. We then propose to transmit the RS with sequential scanning, using a pre-designed codebook with narrow and/or wide beams to approximate the desirable patterns. The proposed design allows flexible choices of the codebook sizes and the associated beam widths to better approximate the desirable patterns. Numerical results demonstrate the effectiveness of the proposed method.

cs.IT

Capacity and Stable Scheduling in Heterogeneous Wireless Networks

Heterogeneous wireless networks (HetNets) provide a means to increase network capacity by introducing small cells and adopting a layered architecture. HetNets allocate resources flexibly through time sharing and cell range expansion/contraction allowing a wide range of possible schedulers. In this paper we define the capacity of a HetNet down link in terms of the maximum number of downloads per second which can be achieved for a given offered traffic density. Given this definition we show that the capacity is determined via the solution to a continuous linear program (LP). If the solution is smaller than 1 then there is a scheduler such that the number of mobiles in the network has ergodic properties with finite mean waiting time. If the solution is greater than 1 then no such scheduler exists. The above results continue to hold if a more general class of schedulers is considered.

cs.NI