SearcharxivSearch

arXiv subjects

Shubin Zhang

Publications and source records attributed to Shubin Zhang.

3 recordsLinked to original sources

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of diffusion transformer blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a novel instruction-guided audio editing framework based on rectified flow matching (RFM), named RFM-Editing 2, built on a hybrid two-stage diffusion transformer. The proposed model performs joint attention over audio and text tokens to establish coarse semantic alignment at the low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at the high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency.

cs.SD

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest, thus demanding precise localization and faithful editing according to the text prompt. Existing training-based and zero-shot methods that rely on full-caption or costly optimization often struggle with complex editing or lack practicality. In this work, we propose a novel end-to-end efficient rectified flow matching-based diffusion framework for audio editing, and construct a dataset featuring overlapping multi-event audio to support training and benchmarking in complex scenarios. Experiments show that our model achieves faithful semantic alignment without requiring auxiliary captions or masks, while maintaining competitive editing quality across metrics.

cs.SD

Resonant multiple-phonon absorption causes efficient anti-Stokes photoluminescence in CsPbBr$_3$ nanocrystals

Lead-halide perovskite nanocrystals such as CsPbBr$_3$, exhibit efficient photoluminescence (PL) up-conversion, also referred to as anti-Stokes photoluminescence (ASPL). This is a phenomenon where irradiating nanocrystals up to 100 meV below gap results in higher energy band edge emission. Most surprising is that ASPL efficiencies approach unity and involve single photon interactions with multiple phonons. This is unexpected given the statistically disfavored nature of multiple-phonon absorption. Here, we report and rationalize near-unity anti-Stokes photoluminescence efficiencies in CsPbBr$_3$ nanocrystals and attribute it to resonant multiple-phonon absorption by polarons. The theory explains paradoxically large efficiencies for intrinsically disfavored, multiple-phonon-assisted ASPL in nanocrystals. Moreover, the developed microscopic mechanism has immediate and important implications for applications of ASPL towards condensed phase optical refrigeration.

cond-mat.mes-hall