SearcharxivSearch

arXiv subjects

Ying Shi

Publications and source records attributed to Ying Shi.

At least 19 recordsLinked to original sources

Training-Free Multi-Step Inference for Target Speaker Extraction

Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder architectures with one-step inference. Inspired by test-time scaling, we propose a training-free multi-step inference method that enables iterative refinement with a frozen pretrained model. At each step, new candidates are generated by interpolating the original mixture and the previous estimate, and the best candidate is selected for further refinement until convergence. Experiments show that, when ground-truth target speech is available, optimizing an intrusive metric (SI-SDRi) yields consistent gains across multiple evaluation metrics. Without ground truth, optimizing non-intrusive metrics (UTMOS or SpkSim) improves the corresponding metric but may hurt others. We therefore introduce joint metric optimization to balance these objectives, enabling controllable extraction preferences for practical deployment.

cs.SD

Real-Time-Capable Betatron Tune Measurement from Schottky Spectra Using Deep Learning and Uncertainty-Aware Kalman Filtering

Betatron tune measurement is essential for beam control in compact proton-therapy synchrotrons, yet conventional peak-detection techniques are not robust under the low signal-to-noise ratio (SNR) conditions typical of these machines. This work presents a lightweight convolutional neural network that performs real-time tune extraction from Schottky spectra with sub-millisecond inference latency and calibrated uncertainty estimates. The model uses attention-based pooling for reliable peak localization and a dual-branch architecture that jointly predicts the tune and its associated uncertainty. Trained with a Laplace negative log-likelihood loss, it produces uncertainty estimates whose magnitude tracks the instantaneous prediction error, which enables uncertainty-aware Kalman filtering for temporal smoothing. Experiments on a large synthetic dataset spanning SNR levels from 0 to $-20$\,dB demonstrate substantial performance gains over traditional peak-detection baselines, while the Kalman filter further suppresses transient outliers in time-series operation. Preliminary validation on operational beam data confirms stable tune tracking without retraining. With only about $2.0\times 10^{4}$ trainable parameters and real-time inference on commodity GPU hardware, the proposed diagnostic offers a practical solution for rapid and accurate betatron tune monitoring in compact medical synchrotrons and similar accelerators.

physics.acc-ph

Mode Control and Dynamic Population Gratings in Quantum-Dot Lasers

Single-mode operation is essential for integrated semiconductor lasers, yet most solutions rely on regrowth, etched gratings, or other complex fabrication steps that limit scalability. We show that quantum-dot (QD) lasers can achieve stable single-mode lasing through a simple cavity design using dynamic population gratings (DPGs). Owing to the low lateral carrier diffusion of QDs, a strong standing-wave-induced carrier grating forms in a reverse-biased saturable absorber and provides self-aligned, mode-selective feedback not attainable in quantum-well devices. A single-ring laser achieves 46 dB side-mode suppression ratio (SMSR), while a dual-ring Vernier laser delivers ($>$ 46 nm) tuning range and up to 52.6 dB SMSR, with continuous-wave operation up to $80\,^{\circ}\mathrm{C}$. The laser remains single-mode under $-10.6$ dB external optical feedback and supports isolator-free data transmission at 32 Gbps. These results establish DPG-enabled QD lasers as a simple and scalable route to tunable, feedback-resilient on-chip light sources for communication, sensing, and reconfigurable photonic systems.

physics.optics

Alzheimers Disease Progression Prediction Based on Manifold Mapping of Irregularly Sampled Longitudinal Data

The uncertainty of clinical examinations frequently leads to irregular observation intervals in longitudinal imaging data, posing challenges for modeling disease progression.Most existing imaging-based disease prediction models operate in Euclidean space, which assumes a flat representation of data and fails to fully capture the intrinsic continuity and nonlinear geometric structure of irregularly sampled longitudinal images. To address the challenge of modeling Alzheimers disease (AD) progression from irregularly sampled longitudinal structural Magnetic Resonance Imaging (sMRI) data, we propose a Riemannian manifold mapping, a Time-aware manifold Neural ordinary differential equation, and an Attention-based riemannian Gated recurrent unit (R-TNAG) framework. Our approach first projects features extracted from high-dimensional sMRI into a manifold space to preserve the intrinsic geometry of disease progression. On this representation, a time-aware Neural Ordinary Differential Equation (TNODE) models the continuous evolution of latent states between observations, while an Attention-based Riemannian Gated Recurrent Unit (ARGRU) adaptively integrates historical and current information to handle irregular intervals. This joint design improves temporal consistency and yields robust AD trajectory prediction under irregular sampling.Experimental results demonstrate that the proposed method consistently outperforms state-of-the-art models in both disease status prediction and cognitive score regression. Ablation studies verify the contributions of each module, highlighting their complementary roles in enhancing predictive accuracy. Moreover, the model exhibits stable performance across varying sequence lengths and missing data rates, indicating strong temporal generalizability. Cross-dataset validation further confirms its robustness and applicability in diverse clinical settings.

cs.CV

MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech

Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. While effective in clean conditions, most existing approaches struggle with mixed keyword spotting--detecting multiple overlapping keywords within a single utterance--a capability essential for real-world applications. We have previously proposed a pre-training approach based on Mix-Training (MT) to tackle the mixed keyword detection problem and demonstrated its efficiency. However, this approach is fully supervised, unable to utilize vast unlabeled data. To this end, we propose Mix-Training HuBERT (MT-HuBERT), a self-supervised learning (SSL) pre-training framework that implements the MT criterion during pre-training. MT-HuBERT predicts, in a self-supervised manner, the clean acoustic units of each constituent signal from contextual cues, in contrast to predicting compositional patterns of mixed speech. Experiments conducted on the Google Speech Commands (GSC v2) corpus demonstrate that our proposed MT-HuBERT consistently outperforms several state-of-the-art baselines in few-shot KWS tasks under both mixed and clean conditions.

cs.SD

Knowledge-Decoupled Functionally Invariant Path with Synthetic Personal Data for Personalized ASR

Fine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic personal data without forgetting real knowledge, and in adapting to personal data without forgetting generic knowledge. Considering that the functionally invariant path (FIP) framework enables model adaptation while preserving prior knowledge, in this letter, we introduce FIP into synthetic-data-augmented personalized ASR models. However, the model still struggles to balance the learning of synthetic, personalized, and generic knowledge when applying FIP to train the model on all three types of data simultaneously. To decouple this learning process and further address the above two challenges, we integrate a gated parameter-isolation strategy into FIP and propose a knowledge-decoupled functionally invariant path (KDFIP) framework, which stores generic and personalized knowledge in separate modules and applies FIP to them sequentially. Specifically, KDFIP adapts the personalized module to synthetic and real personal data and the generic module to generic data. Both modules are updated along personalization-invariant paths, and their outputs are dynamically fused through a gating mechanism. With augmented synthetic data, KDFIP achieves a 29.38% relative character error rate reduction on target speakers and maintains comparable generalization performance to the unadapted ASR baseline.

cs.SD

Yet Another Watermark for Large Language Models

Existing watermarking methods for large language models (LLMs) mainly embed watermark by adjusting the token sampling prediction or post-processing, lacking intrinsic coupling with LLMs, which may significantly reduce the semantic quality of the generated marked texts. Traditional watermarking methods based on training or fine-tuning may be extendable to LLMs. However, most of them are limited to the white-box scenario, or very time-consuming due to the massive parameters of LLMs. In this paper, we present a new watermarking framework for LLMs, where the watermark is embedded into the LLM by manipulating the internal parameters of the LLM, and can be extracted from the generated text without accessing the LLM. Comparing with related methods, the proposed method entangles the watermark with the intrinsic parameters of the LLM, which better balances the robustness and imperceptibility of the watermark. Moreover, the proposed method enables us to extract the watermark under the black-box scenario, which is computationally efficient for use. Experimental results have also verified the feasibility, superiority and practicality. This work provides a new perspective different from mainstream works, which may shed light on future research.

cs.CR

Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling

Recently, cross-attention-based contextual automatic speech recognition (ASR) models have made notable advancements in recognizing personalized biasing phrases. However, the effectiveness of cross-attention is affected by variations in biasing information volume, especially when the length of the biasing list increases significantly. We find that, regardless of the length of the biasing list, only a limited amount of biasing information is most relevant to a specific ASR intermediate representation. Therefore, by identifying and integrating the most relevant biasing information rather than the entire biasing list, we can alleviate the effects of variations in biasing information volume for contextual ASR. To this end, we propose a purified semantic correlation joint modeling (PSC-Joint) approach. In PSC-Joint, we define and calculate three semantic correlations between the ASR intermediate representations and biasing information from coarse to fine: list-level, phrase-level, and token-level. Then, the three correlations are jointly modeled to produce their intersection, so that the most relevant biasing information across various granularities is highlighted and integrated for contextual recognition. In addition, to reduce the computational cost introduced by the joint modeling of three semantic correlations, we also propose a purification mechanism based on a grouped-and-competitive strategy to filter out irrelevant biasing phrases. Compared with baselines, our PSC-Joint approach achieves average relative F1 score improvements of up to 21.34% on AISHELL-1 and 28.46% on KeSpeech, across biasing lists of varying lengths.

cs.CL

Bizard: A Community-Driven Platform for Accelerating and Enhancing Biomedical Data Visualization

Biomedical research increasingly relies on heterogeneous, high-dimensional datasets, yet effective visualization remains hindered by fragmented code resources, steep programming barriers, and limited domain-specific guidance. Bizard is an open-source visualization code repository engineered to streamline data analysis in biomedical research. It aggregates a diverse array of executable visualization scripts, empowering researchers to select and tailor optimal graphical methods for their specific investigative demands. The platform features an intuitive interface equipped with sophisticated browsing and filtering capabilities, exhaustive tutorials, and interactive discussion forums that foster knowledge dissemination. Through its community-driven paradigm, Bizard promotes continual refinement and functional expansion, establishing itself as an essential resource for elevating biomedical data visualization and analytical standards. By harnessing Bizard's infrastructure, researchers can augment their visualization proficiency, propel methodological progress, and enhance interpretive rigor, ultimately accelerating precision medicine and personalized therapeutics. Bizard is freely accessible at https://openbiox.github.io/Bizard/.

q-bio.GN

Intrinsic exciton transport and recombination in single-crystal lead bromide perovskite

Photogenerated carrier transport and recombination in metal halide perovskites are critical to device performance. Despite considerable efforts, sample quality issues and measurement techniques have limited the access to their intrinsic physics. Here, by utilizing high-purity CsPbBr3 single crystals and contact-free transient grating spectroscopy, we directly monitor exciton diffusive transport from 26 to 300 K. As the temperature (T) increases, the carrier mobility ({\mu}) decreases rapidly below 100 K wtih a {\mu}~T^{-3.0} scaling, and then follows a more gradual {\mu}~T^{-1.7} trend at higher temperatures. First-principles calculations perfectly reproduce this experimental trend and reveal that optical phonon scattering governs carrier mobility shifts over the entire temperature range, with a single longitudinal optical mode dominating room-temperature transport. Time-resolved photoluminescence further identifies a substantial increase in exciton radiative lifetime with temperature, attributed to increased exciton population in momentum-dark states caused by phonon scattering. Our findings unambiguously resolve previous theory-experiment discrepancies, providing benchmarks for future optoelectronic design.

cond-mat.mtrl-sci

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing.

cs.CL

High-Accuracy Schottky Diagnostics for Low-SNR Betatron Tune Measurement in Ramping Synchrotrons

This study introduces a novel real-time betatron tune measurement algorithm, utilizing Schottky signals and an FPGA-based backend architecture, specifically designed for rapidly ramping synchrotrons, with particular application to the Shanghai Advanced Proton Therapy (SAPT) facility. The developed algorithm demonstrates improved measurement accuracy under challenging operational conditions, especially in scenarios with limited sampling time and signal-to-noise ratios (SNR) as low as \(-20\) dB. By applying Short-Time Fourier Transform (STFT) analysis, the algorithm effectively accommodates the rapid increase in revolution frequency from 4 MHz to 7.5 MHz over 0.35 seconds, along with tune shifts. A macro-particle simulation methodology is employed to generate Schottky signals, which are then combined with real noise collected from an analog-to-digital converter (ADC) to simulate practical conditions. The proposed betatron tune measurement algorithm integrates advanced spectral processing techniques and an enhanced peak detection algorithm specifically tailored for low SNR conditions. Experimental validation confirms the superior performance of the proposed algorithm over conventional approaches in terms of measurement accuracy, stability, and system robustness, while meeting the stringent operational requirements of proton therapy applications. This innovative approach effectively addresses critical limitations associated with Schottky diagnostics for betatron tune measurement in rapidly ramping synchrotrons operating under low SNR conditions, laying a robust foundation and providing a viable solution for advanced applications in proton therapy and related accelerator physics fields.

physics.acc-ph

Unraveling Intertwined Orders in the Strongly Correlated Kagome Metal CsCr3Sb5

While correlated phenomena of flat bands have been extensively studied in twisted systems, the ordered states that emerge from interactions in the intrinsic flat bands of kagome lattice materials remain largely unexplored. The newly discovered kagome metal CsCr3Sb5 offers a unique and rich platform for this research, as its multi-orbital flat bands at the Fermi surface result in a complex interplay of pressurized superconductivity, antiferromagnetism, a structural phase transition, and density wave orders. Here, using ultrafast optical techniques, we provide strong spectroscopic evidence for a charge density wave transition in CsCr3Sb5, resolving previous ambiguities. Crucially, we identify rotational symmetry breaking that manifests as a three-state Potts-type nematicity. Our elastoresistance measurements directly demonstrate the electronic origin of this order, as the rotational-symmetry-breaking E2g component of the elastoresistance shows a divergent behaviour around the transition temperature. This exotic nematicity results from the lifting of degeneracy of the multi-orbital flat bands, akin to phenomena seen in certain iron-based superconductors. Our study pioneers the investigation of ultrafast dynamics in flat-band systems at the Fermi surface, offering new insights into the interactions between multiple elementary excitations in strongly correlated systems.

cond-mat.str-el

Few-Shot Keyword Spotting from Mixed Speech

Few-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples. A commonly used approach is the pre-training and fine-tuning framework. While effective in clean conditions, this approach struggles with mixed keyword spotting -- simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications. Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario. In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech. Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase. Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions.

cs.SD

Serialized Output Training by Learned Dominance

Serialized Output Training (SOT) has showcased state-of-the-art performance in multi-talker speech recognition by sequentially decoding the speech of individual speakers. To address the challenging label-permutation issue, prior methods have relied on either the Permutation Invariant Training (PIT) or the time-based First-In-First-Out (FIFO) rule. This study presents a model-based serialization strategy that incorporates an auxiliary module into the Attention Encoder-Decoder architecture, autonomously identifying the crucial factors to order the output sequence of the speech components in multi-talker speech. Experiments conducted on the LibriSpeech and LibriMix databases reveal that our approach significantly outperforms the PIT and FIFO baselines in both 2-mix and 3-mix scenarios. Further analysis shows that the serialization module identifies dominant speech components in a mixture by factors including loudness and gender, and orders speech components based on the dominance score.

cs.SD

A Glance is Enough: Extract Target Sentence By Looking at A keyword

This paper investigates the possibility of extracting a target sentence from multi-talker speech using only a keyword as input. For example, in social security applications, the keyword might be "help", and the goal is to identify what the person who called for help is articulating while ignoring other speakers. To address this problem, we propose using the Transformer architecture to embed both the keyword and the speech utterance and then rely on the cross-attention mechanism to select the correct content from the concatenated or overlapping speech. Experimental results on Librispeech demonstrate that our proposed method can effectively extract target sentences from very noisy and mixed speech (SNR=-3dB), achieving a phone error rate (PER) of 26\%, compared to the baseline system's PER of 96%.

cs.CL

Matrix-valued $\theta$-deformed bi-orthogonal polynomials, Non-commutative Toda theory and B\"acklund transformation

This paper is devoted to revealing the relationship between matrix-valued $\theta$-deformed bi-orthogonal polynomials and non-commutative Toda-type hierarchies. In this procedure, Wronski quasi-determinants are widely used and play the role of non-commutative $\tau$-functions. At the same time, B\"acklund transformations are realized by using a moment modification method and non-commutative $\theta$-deformed Volterra hierarchies are obtained, which contain the known examples of the Itoh-Narita-Bogoyavlensky lattices and the fractional Volterra hierarchy.

nlin.SI

Spot keywords from very noisy and mixed speech

Most existing keyword spotting research focuses on conditions with slight or moderate noise. In this paper, we try to tackle a more challenging task: detecting keywords buried under strong interfering speech (10 times higher than the keyword in amplitude), and even worse, mixed with other keywords. We propose a novel Mix Training (MT) strategy that encourages the model to discover low-energy keywords from noisy and mixed speech. Experiments were conducted with a vanilla CNN and two EfficientNet (B0/B2) architectures. The results evaluated with the Google Speech Command dataset demonstrated that the proposed mix training approach is highly effective and outperforms standard data augmentation and mixup training.

cs.SD