SearcharxivSearch

arXiv subjects

Wei Rao

Publications and source records attributed to Wei Rao.

At least 19 recordsLinked to original sources

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall performance. For instance, applying AGC before SE may inadvertently amplify background noise, while prioritizing SE tends to over-suppress low-volume speech. To address these limitations, we propose SE-AGCNet, an end-to-end framework that jointly optimizes SE and AGC. Tailored for meeting scenarios with significant volume variations, SE-AGCNet leverages the synergy between the two tasks: SE preserves quiet speech, thereby facilitating effective volume adjustment by the AGC component. Furthermore, we propose a specialized data simulation pipeline, SE-AGC-DataGen, and incorporate standardized loudness evaluation metrics: integrated loudness (LUFS), short-term loudness (St LUFS), and LRA. Experiments show that SE-AGCNet consistently achieves target loudness while improving speech quality and ASR accuracy over competitive baselines.

eess.AS

VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations

With the rapid deployment of speech generation systems in open environments, providing verifiable source attribution and copyright accountability for audio content has become critical. A gap in current research is the lack of a unified benchmark that systematically compares different watermark injection methods under realistic distribution shifts. To address this, we build VoxWatermark by applying 10 watermarking methods (4 neural and 6 traditional) with unified injection and annotation on multilingual, multi-source corpora, and introducing no-box, black-box, and white-box perturbations to simulate real recording and transmission conditions. Based on this benchmark, we propose AudioWMD as a robust baseline detector for large-scale, multi-method, cross-distribution settings. Results show that injection-method diversity and distribution shifts affect detection stability, while validating the effectiveness and scalability of AudioWMD. Dataset and code are publicly available.

eess.AS

Injectable Thermochemical Micro-Explosion for Prompt Thrombolysis via Liquid Alkali Metal

Thrombotic vascular diseases contribute to significant global mortality, yet current therapeutic strategies face persistent challenges including bleeding risks, suboptimal efficiency, and procedural complexity. Here, we report a micro-explosive thermochemical thrombolysis (METCT) therapy via injectable liquid alkali metal (LAM) encapsulated in dimethyl silicone (LAM@oil), which enables prompt, efficient and safe vascular recanalization within an ultrafast timeframe (< 90 seconds). This LAM@oil system effectively disrupts thrombus tissue through a synergistic triple-action mechanism: Mechanical micro-explosions forces, alkaline ablation due to highly localized exothermic chemical reactions, and thermal thrombolysis mediated by elevated temperature. Upon thrombolysis completion, the non-toxic reaction byproducts (sodium and potassium ions) exhibit physiologically biocompatible and metabolizable effects. Critically, the LAM@oil demonstrates significantly higher thrombolytic efficacy compared to clinically available thrombolytic drugs (residual thrombus area percent 10.87%+-7.16% for LAM@oil vs. 80.86%+-13.32% for urokinase), with no associated bleeding risks. This strategy opens a byproduct-free, cost-effective, and high-efficiency alternative to conventional thrombolytics, holding big potential for clinical translation in acute thrombosis management.

physics.med-ph

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework that explicitly anchors each reasoning step to specific visual evidence regions, enabling compositional and multi-step decision-making. Formally, Chain-of-Glimpse formulates video reasoning as a step-by-step process that incrementally builds spatially grounded traces around task-relevant visual objects, thereby mitigating over-reliance on saliency-driven cues. Specifically, Chain-of-Glimpse features a search-guided controller, optimized via reinforcement learning with a format reward that significantly incentivizes grounding capability, to iteratively ground visual evidence regions and form reliable reasoning trajectories, yielding accurate and interpretable multi-step decisions. Extensive evaluations on both in domain NExTQA and out-of-domain Video-Holmes, CG-Bench Reasoning, and VRBench benchmarks demonstrate consistent performance gains, robustness and generalization of Chain-of-Glimpse across diverse video reasoning tasks.

cs.CV

A note on piercing discrete rectangles

In 2008, Halman proved a discrete Helly-type theorem for axis-parallel boxes in $\mathbb R^d$. Very recently, this result was extended to the $(p,q)$ setting with $p \geq q \geq d+1$ by Edwards and Sober\'on, and subsequently to the case $p \geq q \geq 2$ by Gangopadhyay, Polyanskii, and the author of this paper. In this paper, we obtain improved bounds for the $(p,q)$ problem in the case $q=2$ and $d=2$. More precisely, our main result asserts that for any integer $p \geq 2$, any set $P \subseteq \mathbb R^2$, and any finite family $\mathcal B$ of axis-parallel rectangles in $\mathbb R^2$ such that every rectangle contains a point of $P$, if among every $p$ rectangles there exist two whose intersection contains a point of $P$, then there exists a subset $S \subseteq P$ of size at most $O\!\bigl( (p \log \log p)^2 \bigr)$ such that every rectangle contains a point of $S$. Moreover, when $p=2$, the size of $S$ can be bounded by $4$.

math.CO

Training-Free Intelligibility-Guided Observation Addition for Noisy ASR

Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce artifacts that harm recognition. Observation addition (OA) addressed this issue by fusing noisy and SE enhanced speech, improving recognition without modifying the parameters of the SE or ASR models. This paper proposes an intelligibility-guided OA method, where fusion weights are derived from intelligibility estimates obtained directly from the backend ASR. Unlike prior OA methods based on trained neural predictors, the proposed method is training-free, reducing complexity and enhances generalization. Extensive experiments across diverse SE-ASR combinations and datasets demonstrate strong robustness and improvements over existing OA baselines. Additional analyses of intelligibility-guided switching-based alternatives and frame versus utterance-level OA further validate the proposed design.

eess.AS

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenTSE, a two-stage decoder-only generative LM for TSE: Stage-1 predicts coarse semantic tokens, and Stage-2 generates fine acoustic tokens. Separating semantics and acoustics stabilizes decoding and yields more accurate target speech. Both stages use continuous SSL or codec embeddings, offering richer context than discretized-prompt methods. To reduce exposure bias, we employ a Frozen-LM Conditioning training strategy that conditions the LMs on predicted tokens from earlier checkpoints to reduce the gap between teacher-forcing training and autoregressive inference. We further apply DPO to better align outputs with perceptual preferences. Experiments on Libri2Mix show that GenTSE surpasses previous LM-based systems in speech quality, intelligibility, and speaker consistency.

eess.AS

New Helly-type results for discrete boxes: Quantitative colorful and $(p,q)$-variants

In 2008, Halman showed that for any finite set $P\subset \mathbb R^d$ and any finite family $\mathcal{B}$ of axis-parallel boxes in $\mathbb{R}^d$, if the intersection of $P$ and any subfamily $\mathcal{B}' \subseteq\mathcal{B}$ of size at most $2d$ is non-empty, then the intersection of $P$ and $\mathcal{B}$ is also non-empty. Very recently Edwards and Sober\'on initiated the study of quantitative colorful version for $2d$ families, $(p,q)$-type variation for $p\geq q\geq d+1$, and other extensions of this Helly-type result by Halman. In this paper, we study the quantitative colorful Halman problem for $2d-1$ families as well its $(p,q)$-type variation for $p\geq q\geq 2$. Specifically, our main result asserts that for any finite set $P$ and finite families of boxes $\mathcal{B}_1,\dots,\mathcal{B}_{2d-1}$ in $\mathbb R^d$, where $d\geq 2$, if every transversal $\mathcal{B}$ for the families has an intersection $\bigcap \mathcal{B}$ containing at least $n$ points of $P$, then there exist $j\in[2d-1]$ and a subset of $P$ of size at most \[ 2n+\Big\lfloor \frac{n-1}{d \cdot 2^{d-1}} \Big\rfloor, \] such that each box of $\mathcal{B}_j$ contains at least $n$ points of this subset.

math.CO

Twist Bilayer Photonic slab's Angle-DependentGuided Resonance Analysis based on Multiple Scattering

We present an analysis of the transmission spectra of the twisted bilayer photonic slabs using a modified rigorous coupled wave (RCWA) analysis, where the evanescent bases are replaced by bases with non-zero flux density. By utilizing the modified RCWA we demonstrate the calculation of eigenmodes, which has not been realized before. To counter for the transmission property, we propose a five-layer uniform slab approximation, with an accuracy around 0.04a/c, which is more straightforward and accessible for optical engineers compared to work by Lou et al. [Phys. Rev. Lett. 126, 136101]. The moir\'e pattern perturbation induces a split of resonance, which show great potential for engineering the band structure. Moreover, We observe two distinct transmission phases: the angle-dependent phase and Fabry-P\'erot phase, which is explained by a coupled-mode theory (CMT) with expanded channels brought by the modified eigenmodes. Our work provides a theoretical framework for the design and optimization of twisted bilayer photonic devices.

physics.optics

Helly-type theorems for separated $d$-intervals

A separated $d$-interval is defined as a disjoint union of $d$ convex sets from the real line $\mathbb R$. In this paper, we establish a series of Helly-type theorems for convexity spaces derived from separated $d$-intervals. Our results encompass the Radon number, Helly number, colorful Helly number, fractional Helly number, colorful fractional Helly theorem, $(p,q)$ theorem, and two kinds of colorful $(p,q)$ theorems for these convexity spaces. The primary tools employed in our proofs involve simplicial complexes and collapsibility.

math.CO

Liquid Metal Printed Superconducting Circuits

Since the discovery of superconductor one hundred years ago, tremendous theoretical and technological progresses have been achieved. The zero resistance and complete diamagnetism of superconducting materials promise many possibilities in diverse fields. However, the complexity and expensive manufacturing costs associated with the time-consuming superconductor fabrication process may retard their practices in a large extent. Here, via liquid metal printing we proposed to quickly fabricate superconducting electronics which can work at the prescribed cryogenic temperatures. By way of the room temperature fluidity of liquid metal composite inks, such one-step printing allows to pattern various superconducting circuits on the desired substrate. As the first-ever conceptual trial, the most easily available gallium-based liquid alloy inks were particularly adopted to composite with copper particles to achieve superconductivity under specific temperatures around 6.4K. Further, a series of liquid metal alloy and particles loaded composites were screened out and comparatively interpreted regarding their superconducting properties and potential values as printable inks in fabricating superconducting devices. The cost-effective feature and straightforward adaptability of the fabrication principle were evaluated. This work suggests an easy-going way for fabricating ending user superconducting devices, which may warrant more promising investigations and practices in the coming time.

cond-mat.supr-con

Hierarchical speaker representation for target speaker extraction

Target speaker extraction aims to isolate a specific speaker's voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the voice of the target speaker. However, the representation of the speaker embedding is too simplistic, often being merely a 1*1024 vector. This dense information makes it difficult for the separation network to harness effectively. To address this limitation, we introduce a pioneering methodology called Hierarchical Representation (HR) that seamlessly fuses anchor data across granular and overarching 5 layers of the separation network, enhancing the precision of target extraction. HR amplifies the efficacy of anchors to improve target speaker isolation. On the Libri-2talker dataset, HR substantially outperforms state-of-the-art time-frequency domain techniques. Further demonstrating HR's capabilities, we achieved first place in the prestigious ICASSP 2023 Deep Noise Suppression Challenge. The proposed HR methodology shows great promise for advancing target speaker extraction through enhanced anchor utilization.

cs.SD

Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization

The scarcity of labeled audio-visual datasets is a constraint for training superior audio-visual speaker diarization systems. To improve the performance of audio-visual speaker diarization, we leverage pre-trained supervised and self-supervised speech models for audio-visual speaker diarization. Specifically, we adopt supervised~(ResNet and ECAPA-TDNN) and self-supervised pre-trained models~(WavLM and HuBERT) as the speaker and audio embedding extractors in an end-to-end audio-visual speaker diarization~(AVSD) system. Then we explore the effectiveness of different frameworks, including Transformer, Conformer, and cross-attention mechanism, in the audio-visual decoder. To mitigate the degradation of performance caused by separate training, we jointly train the audio encoder, speaker encoder, and audio-visual decoder in the AVSD system. Experiments on the MISP dataset demonstrate that the proposed method achieves superior performance and obtained third place in MISP Challenge 2022.

eess.AS

The FlySpeech Audio-Visual Speaker Diarization System for MISP Challenge 2022

This paper describes the FlySpeech speaker diarization system submitted to the second \textbf{M}ultimodal \textbf{I}nformation Based \textbf{S}peech \textbf{P}rocessing~(\textbf{MISP}) Challenge held in ICASSP 2022. We develop an end-to-end audio-visual speaker diarization~(AVSD) system, which consists of a lip encoder, a speaker encoder, and an audio-visual decoder. Specifically, to mitigate the degradation of diarization performance caused by separate training, we jointly train the speaker encoder and the audio-visual decoder. In addition, we leverage the large-data pretrained speaker extractor to initialize the speaker encoder.

cs.SD

MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation

The previous SpEx+ has yielded outstanding performance in speaker extraction and attracted much attention. However, it still encounters inadequate utilization of multi-scale information and speaker embedding. To this end, this paper proposes a new effective speaker extraction system with multi-scale interfusion and conditional speaker modulation (ConSM), which is called MC-SpEx. First of all, we design the weight-share multi-scale fusers (ScaleFusers) for efficiently leveraging multi-scale information as well as ensuring consistency of the model's feature space. Then, to consider different scale information while generating masks, the multi-scale interactive mask generator (ScaleInterMG) is presented. Moreover, we introduce ConSM module to fully exploit speaker embedding in the speech extractor. Experimental results on the Libri2Mix dataset demonstrate the effectiveness of our improvements and the state-of-the-art performance of our proposed MC-SpEx.

cs.SD

Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction

This paper describes a real-time General Speech Reconstruction (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This novel proposed system is a two-stage architecture, in which the speech restoration is performed, and then cascaded by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is performed with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2.

cs.SD

A photo-mechanical coupling theory for photoisomerization hydrogel considering the distribution state of molecular chains

Owing to the possibility of controlling its specific mechanical behaviors up taken by irradiated by light at particular wavelengths, the photoisomerization hydrogels have a broad range of potential applications. A theory connecting the optical excitation to mechanical behavior is essential to precisely control the photo-mechanical behaviors of the hydrogel. In this work, a photo-mechanical coupling theory is developed to describe the photo-mechanical responses of photoisomerization hydrogels within the framework of finite deformation continuum thermodynamics. In the model, the deformation gradient is decomposed into two parts to effectively model the light induced deformation and the elastic one. To consider the effect of the optical excitation on mechanical behaviors, we first investigate the transporting mechanism of light in hydrogel, as well as the photochemical reaction process; and we then explore the disturbance of light irradiation on the equilibrium of the thermodynamic system of hydrogel, as well as the relationship of conformational entropy of hydrogel network with the photochemical reaction; finally, based on the entropy elasticity theory, we propose a new free energy function of the photosensitive hydrogel to consider the effect of molecular chain distribution evolution on the stiffness of the hydrogel network. With the implementation of the proposed model, we study the photo-mechanical behaviors and mechanical properties of photoisomerization hydrogels. The present research is helpful for understanding the multi-field coupling behaviors of the photosensitive hydrogel, and then providing guidelines for the application of photoisomerization hydrogel.

physics.chem-ph

Inter-SubNet: Speech Enhancement with Subband Interaction

Subband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral information has a negative impact on the performance of these subband-based approaches. To this end, this paper introduces the subband interaction as a new way to complement the subband model with the global spectral information such as cross-band dependencies and global spectral patterns, and proposes a new lightweight single-channel speech enhancement framework called Interactive Subband Network (Inter-SubNet). Experimental results on DNS Challenge - Interspeech 2021 dataset show that the proposed Inter-SubNet yields a significant improvement over the subband model and outperforms other state-of-the-art speech enhancement approaches, which demonstrate the effectiveness of subband interaction.

cs.SD