SearcharxivSearch

arXiv subjects

Yifan Liang

Publications and source records attributed to Yifan Liang.

12 recordsLinked to original sources

Joint Synchronization and Sensing in Networked ISAC via Structured Canonical Polyadic Decomposition

Networked integrated sensing and communication (ISAC) offers significant potential for next-generation wireless systems. By exploiting spatial diversity through the cooperation of multiple base stations (BSs), this architecture expands coverage and achieves enhanced sensing performance. However, accurate sensing in networked ISAC requires time-frequency synchronization among BSs. Existing synchronization methods for networked ISAC suffer from inter-path interference caused by sensing channel compression. To address this problem, this paper proposes a structured canonical polyadic decomposition (SCPD) algorithm that effectively separates the multipath components of the sensing channel. Benefiting from this separation, SCPD achieves joint network-level synchronization and multi-target parameter estimation. We establish theoretical identifiability conditions for SCPD and show that it asymptotically achieves the Cram\'{e}r-Rao bound. Furthermore, by incorporating parameters estimated from different BS pairs, we propose a multi-target tracking algorithm designed for the continuous operation of the system. The proposed algorithm tracks both the trajectories and velocities of moving targets by leveraging geometric diversity. Utilizing tracking results from the previous snapshot, an adaptive beamforming scheme is also developed to improve tracking performance in the next snapshot. Simulation results demonstrate that the proposed algorithms achieve superior accuracy and outlier robustness for both synchronization and sensing in networked ISAC, outperforming traditional approaches.

eess.SP

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in this task remains largely unexplored. In this paper, we introduce SLD-L2S, a novel L2S framework built upon a hierarchical subspace latent diffusion model. Our method aims to directly map visual lip movements to the continuous latent space of a pre-trained neural audio codec, thereby avoiding the information loss inherent in traditional intermediate representations. The core of our method is a hierarchical architecture that processes visual representations through multiple parallel subspaces, initiated by a subspace decomposition module. To efficiently enhance interactions within and between these subspaces, we design the diffusion convolution block (DiCB) as our network backbone. Furthermore, we employ a reparameterized flow matching technique to directly generate the target latent vectors. This enables a principled inclusion of speech language model (SLM) and semantic losses during training, moving beyond conventional flow matching objectives and improving synthesized speech quality. Our experiments show that SLD-L2S achieves state-of-the-art generation quality on multiple benchmark datasets, surpassing existing methods in both objective and subjective evaluations.

eess.AS

GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does SNR fail in measuring audio quality? And how to improve its reliability as an objective metric? In this paper, we identify the inadequate measurement of phase distance as a pivotal factor and propose to reformulate SNR with specially designed phase-distance terms, yielding an improved metric named GOMPSNR. We further extend the newly proposed formulation to derive two novel categories of loss function, corresponding to magnitude-guided phase refinement and joint magnitude-phase optimization, respectively. Besides, extensive experiments are conducted for an optimal combination of different loss functions. Experimental results on advanced neural vocoders demonstrate that our proposed GOMPSNR exhibits more reliable error measurement than SNR. Meanwhile, our proposed loss functions yield substantial improvements in model performance, and our wellchosen combination of different loss functions further optimizes the overall model capability.

cs.SD

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To address this issue, we propose Lexical Tone-Aware Lip-to-Speech (LTA-L2S). To tackle viseme-to-phoneme complexity, our model adapts an English pre-trained audio-visual self-supervised learning (SSL) model via a cross-lingual transfer learning strategy. This strategy not only transfers universal knowledge learned from extensive English data to the Mandarin domain but also circumvents the prohibitive cost of training such a model from scratch. To specifically model lexical tones and enhance intelligibility, we further employ a flow-matching model to generate the F0 contour. This generation process is guided by ASR-fine-tuned SSL speech units, which contain crucial suprasegmental information. The overall speech quality is then elevated through a two-stage training paradigm, where a flow-matching postnet refines the coarse spectrogram from the first stage. Extensive experiments demonstrate that LTA-L2S significantly outperforms existing methods in both speech intelligibility and tonal accuracy.

cs.SD

OpenHAIV: A Framework Towards Practical Open-World Learning

Substantial progress has been made in various techniques for open-world recognition. Out-of-distribution (OOD) detection methods can effectively distinguish between known and unknown classes in the data, while incremental learning enables continuous model knowledge updates. However, in open-world scenarios, these approaches still face limitations. Relying solely on OOD detection does not facilitate knowledge updates in the model, and incremental fine-tuning typically requires supervised conditions, which significantly deviate from open-world settings. To address these challenges, this paper proposes OpenHAIV, a novel framework that integrates OOD detection, new class discovery, and incremental continual fine-tuning into a unified pipeline. This framework allows models to autonomously acquire and update knowledge in open-world environments. The proposed framework is available at https://haiv-lab.github.io/openhaiv .

cs.CV

OpenEarthSensing: Large-Scale Fine-Grained Benchmark for Open-World Remote Sensing

The advancement of remote sensing, including satellite systems, facilitates the continuous acquisition of remote sensing imagery globally, introducing novel challenges for achieving open-world tasks. Deployed models need to continuously adjust to a constant influx of new data, which frequently exhibits diverse shifts from the data encountered during the training phase. To effectively handle the new data, models are required to detect semantic shifts, adapt to covariate shifts, and continuously update their parameters without forgetting learned knowledge, as has been considered in works on a variety of open-world tasks. However, existing studies are typically conducted within a single dataset to simulate realistic conditions, with a lack of large-scale benchmarks capable of evaluating multiple open-world tasks. In this paper, we introduce \textbf{OpenEarthSensing (OES)}, a large-scale fine-grained benchmark for open-world remote sensing. OES includes 189 scene and object categories, covering the vast majority of potential semantic shifts that may occur in the real world. Additionally, to provide a more comprehensive testbed for evaluating the generalization performance, OES encompasses five data domains with significant covariate shifts, including two RGB satellite domains, one RGB aerial domain, one multispectral RGB domain, and one infrared domain. We evaluate the baselines and existing methods for diverse tasks on OES, demonstrating that it serves as a meaningful and challenging benchmark for open-world remote sensing. The proposed dataset OES is available at https://haiv-lab.github.io/OES.

cs.CV

Demonstration of Direct-amplification Enabled Harmonic Generation in an Ultraviolet Free-Electron Laser

We report the experimental demonstration of direct-amplification enabled harmonic generation in an ultraviolet free-electron laser (FEL) driven by a low-intensity seed laser. By employing a versatile undulator configuration that enables seed amplification and harmonic generation within a unified setup, we achieved over 100-fold energy gain of the seed and observed exponential growth at the second harmonic. The results demonstrate that a sufficiently long modulator can not only amplify a weak seed but also induce strong energy modulation of the electron beam, enabling efficient harmonic bunching. This method markedly relaxes the power requirements on external seed lasers and presents a viable route toward high-repetition-rate, fully coherent FELs

physics.acc-ph

Stability Enhancement of a Self-Amplified Spontaneous Emission Free-electron Laser with Bunching Containment

The self-amplified spontaneous emission (SASE) mechanism, the fundamental operating principle of numerous free-electron laser (FEL) facilities, is driven by electron beam shot noise and leads to significant fluctuations in the output pulse energy. This study presents a robust method for improving pulse energy stability by incorporating a dispersion element that introduces longitudinal dispersion into the electron beam during the exponential growth phase of the SASE process. At this phase, the density modulation of the electron beam, characterized by the bunching factor, undergoes large fluctuations, resulting in substantial variations in the emitted radiation power. The introduction of longitudinal dispersion allows for controlled manipulation of the bunching distribution, suppressing fluctuations and enhancing pulse energy stability. The stabilization mechanism is explained in this paper, and its impact on the radiation properties is analyzed for both the standard SASE scheme and advanced lasing setups, such as a two-stage lasing process for two-color pulse generation, with the initial stage operating in SASE mode.

physics.acc-ph

NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing

Recent advancements in visual speech recognition (VSR) have promoted progress in lip-to-speech synthesis, where pre-trained VSR models enhance the intelligibility of synthesized speech by providing valuable semantic information. The success achieved by cascade frameworks, which combine pseudo-VSR with pseudo-text-to-speech (TTS) or implicitly utilize the transcribed text, highlights the benefits of leveraging VSR models. However, these methods typically rely on mel-spectrograms as an intermediate representation, which may introduce a key bottleneck: the domain gap between synthetic mel-spectrograms, generated from inherently error-prone lip-to-speech mappings, and real mel-spectrograms used to train vocoders. This mismatch inevitably degrades synthesis quality. To bridge this gap, we propose Natural Lip-to-Speech (NaturalL2S), an end-to-end framework integrating acoustic inductive biases with differentiable speech generation components. Specifically, we introduce a fundamental frequency (F0) predictor to capture prosodic variations in synthesized speech. The predicted F0 then drives a Differentiable Digital Signal Processing (DDSP) synthesizer to generate a coarse signal which serves as prior information for subsequent speech synthesis. Additionally, instead of relying on a reference speaker embedding as an auxiliary input, our approach achieves satisfactory performance on speaker similarity without explicitly modelling speaker characteristics. Both objective and subjective evaluation results demonstrate that NaturalL2S can effectively enhance the quality of the synthesized speech when compared to state-of-the-art methods. Our demonstration page is accessible at https://yifan-liang.github.io/NaturalL2S/.

cs.SD

Cascaded high-gradient terahertz-driven acceleration of relativistic electron beams

Terahertz (THz)-driven acceleration has recently emerged as a new route for delivering ultrashort bright electron beams efficiently, reliably, and in a compact setup. Many THz-driven acceleration related working schemes and key technologies have been successfully demonstrated and are continuously being improved to new limits. However, the achieved acceleration gradient and energy gain remain low, and the potential physics and technical challenges in the high field and high energy regime are still under-explored. Here we report a record energy gain of 170 keV in a single-stage configuration, and demonstrate the first cascaded acceleration of a relativistic beam with a 204 keV energy gain in a two-stages setup. Whole-bunch acceleration is accomplished with an average accelerating gradient of 85 MV/m and a peak THz electric field of 1.1 GV/m. This proof-of-principle result is a crucial advance in THz-driven acceleration with a major impact on future electron sources and related scientific discoveries.

physics.acc-ph

Source-Channel Coding and Separation for Generalized Communication Systems

We consider transmission of stationary and ergodic sources over non-ergodic composite channels with channel state information at the receiver (CSIR). Previously we introduced alternate capacity definitions to Shannon capacity, including the capacity versus outage and the expected capacity. These generalized definitions relax the constraint of Shannon capacity that all transmitted information must be decoded at the receiver. In this work alternate end-to-end distortion metrics such as the distortion versus outage and the expected distortion are introduced to relax the constraint that a single distortion level has to be maintained for all channel states. For transmission of stationary and ergodic sources over stationary and ergodic channels, the classical Shannon separation theorem enables separate design of source and channel codes and guarantees optimal performance. For generalized communication systems, we show that different end-to-end distortion metrics lead to different conclusions about separation optimality even for the same source and channel models.

cs.IT

Capacity Definitions for General Channels with Receiver Side Information

We consider three capacity definitions for general channels with channel side information at the receiver, where the channel is modeled as a sequence of finite dimensional conditional distributions not necessarily stationary, ergodic, or information stable. The {\em Shannon capacity} is the highest rate asymptotically achievable with arbitrarily small error probability. The {\em capacity versus outage} is the highest rate asymptotically achievable with a given probability of decoder-recognized outage. The {\em expected capacity} is the highest average rate asymptotically achievable with a single encoder and multiple decoders, where the channel side information determines the decoder in use. As a special case of channel codes for expected rate, the code for capacity versus outage has two decoders: one operates in the non-outage states and decodes all transmitted information, and the other operates in the outage states and decodes nothing. Expected capacity equals Shannon capacity for channels governed by a stationary ergodic random process but is typically greater for general channels. These alternative capacity definitions essentially relax the constraint that all transmitted information must be decoded at the receiver. We derive capacity theorems for these capacity definitions through information density. Numerical examples are provided to demonstrate their connections and differences. We also discuss the implication of these alternative capacity definitions for end-to-end distortion, source-channel coding and separation.

cs.IT