Searcharxiv⌕ Search

arXiv subjects

Ya Li

Publications and source records attributed to Ya Li.

At least 37 records · Page 2Linked to original sources

Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a multi-loss learning (MLL) framework integrating an energy-adaptive mixup (EAM) method and a frame-level attention module (FLAM). The EAM method leverages SNR-based augmentation to generate diverse speech samples capturing subtle emotional variations. FLAM enhances frame-level feature extraction for multi-frame emotional cues. Our MLL strategy combines Kullback-Leibler divergence, focal, center, and supervised contrastive loss to optimize learning, address class imbalance, and improve feature separability. We evaluate our method on four widely used SER datasets: IEMOCAP, MSP-IMPROV, RAVDESS, and SAVEE. The results demonstrate our method achieves state-of-the-art performance, suggesting its effectiveness and robustness.

cs.SD↗

Hello-Chat: Towards Realistic Social Audio Interactions

Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting in a robotic "read-speech" style that lacks the spontaneity and emotional resonance of real human interaction. In this report, we introduce Hello-Chat, an end-to-end audio language model designed for realistic social scenarios. By leveraging a massive dataset of real-life conversations and employing a modality-interleaved training strategy, Hello-Chat achieves a breakthrough in anthropomorphic generation. Experimental results show that our model not only reaches state-of-the-art (SOTA) performance on specific audio understanding tasks but also significantly outperforms existing baselines in prosodic naturalness and emotional alignment, paving the way for the next generation of empathetic AI agents.

cs.SD↗

RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS

Differentiable reinforcement learning (RL) frameworks like DiffRO offer a powerful approach for controllable text-to-speech (TTS), but are vulnerable to reward hacking, particularly for nuanced tasks like emotion control. The policy model can exploit a vanilla Reward Model (RM) by generating acoustic artifacts to achieve spurious rewards, but at the cost of degrading perceptual quality. To address this, we propose Robust Reward Policy Optimization (RRPO), a novel framework that employs a hybrid regularization scheme. This scheme develops a robust RM whose reward signal is more reliably aligned with human perception, compelling the policy to abandon detrimental shortcuts and instead learn the complex features of genuine emotions. Our ablation study confirms the enhanced robustness of our RM, as evidenced by its strong cross-lingual generalization. The subjective evaluation demonstrates that this robust RM effectively mitigates reward hacking, leading to significant improvements in both emotional expressiveness and naturalness over all baselines. Demo page: https://lrwinr.github.io/RRPO-CosyVoice.

cs.SD↗

Secrecy Capacity Analysis and Beamforming Optimization for MIMO-VLC Wiretap Channels

This paper investigates a multiple-input multipleoutput (MIMO) visible light communication (VLC) wiretap channel consisting of a transmitter, a legitimate receiver, and an eavesdropper. The optical input is subject to both peakand average-intensity constraints. By applying the generalized entropy-power inequality to truncated exponential inputs, we derive a novel closed-form expression for the achievable secrecy rate for general MIMO VLC configurations. To enhance transmission confidentiality, a fully-connected beamforming scheme is proposed, along with a low-complexity sub-connected alternative. Although the resulting beamforming design problems are nonconvex, they are efficiently addressed by transforming them into a sequence of convex subproblems solvable via the successive convex approximation framework. Numerical results demonstrate that the proposed schemes achieve significant secrecy performance improvements compared with the benchmark scheme.

cs.IT↗

Cabibbo-suppressed charged-current semileptonic decays of $Ξ_b$ baryons

We present the first perturbative QCD calculations of the $Ξ_b \to (Λ, Σ)$ transition form factors at leading order in $α_s$, which govern the Cabibbo-suppressed semileptonic decays $Ξ_b \to (Λ, Σ)\ell ν_\ell$ with $\ell = e, μ, τ$. Using these form factors, we evaluate differential and integrated branching fractions and angular observables within the helicity formalism. The branching ratios are predicted to be of order $10^{-4}$ for $Σ$ final states and $10^{-5}$ for $Λ$ final states, making them accessible to ongoing experiments such as LHCb. Ratios of decay rates between $τ$ and $e$ channels are also provided, offering new probes of lepton-flavor universality. Lepton-mass effects are found to significantly impact the integrated angular observables. Furthermore, a combined analysis of $b \to u$ and $b \to c$ transitions in $Ξ_b$ decays yields sub-percent precision for the ratios $\mathcal{R}_\ell(Σ/Ξ_c)$, enabling an independent determination of $|V_{ub}/V_{cb}|$ once the relevant decay-rate measurements become available.

hep-ph↗

The Renaissance of Expert Systems: Optical Recognition of Printed Chinese Jianpu Musical Scores with Lyrics

Large-scale optical music recognition (OMR) research has focused mainly on Western staff notation, leaving Chinese Jianpu (numbered notation) and its rich lyric resources underexplored. We present a modular expert-system pipeline that converts printed Jianpu scores with lyrics into machine-readable MusicXML and MIDI, without requiring massive annotated training data. Our approach adopts a top-down expert-system design, leveraging traditional computer-vision techniques (e.g., phrase correlation, skeleton analysis) to capitalize on prior knowledge, while integrating unsupervised deep-learning modules for image feature embeddings. This hybrid strategy strikes a balance between interpretability and accuracy. Evaluated on The Anthology of Chinese Folk Songs, our system massively digitizes (i) a melody-only collection of more than 5,000 songs (> 300,000 notes) and (ii) a curated subset with lyrics comprising over 1,400 songs (> 100,000 notes). The system achieves high-precision recognition on both melody (note-wise F1 = 0.951) and aligned lyrics (character-wise F1 = 0.931).

cs.CV↗

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment requires synchronizing both temporal and semantic information across modalities, while fusion involves integrating these aligned features into a unified representation. Existing methods often address alignment or fusion in isolation, leading to limitations in performance and efficiency. To tackle these issues, we propose a novel framework called Dual-stream Alignment with Hierarchical Bottleneck Fusion (DashFusion). Firstly, dual-stream alignment module synchronizes multimodal features through temporal and semantic alignment. Temporal alignment employs cross-modal attention to establish frame-level correspondences among multimodal sequences. Semantic alignment ensures consistency across the feature space through contrastive learning. Secondly, supervised contrastive learning leverages label information to refine the modality features. Finally, hierarchical bottleneck fusion progressively integrates multimodal information through compressed bottleneck tokens, which achieves a balance between performance and computational efficiency. We evaluate DashFusion on three datasets: CMU-MOSI, CMU-MOSEI, and CH-SIMS. Experimental results demonstrate that DashFusion achieves state-of-the-art performance across various metrics, and ablation studies confirm the effectiveness of our alignment and fusion techniques. The codes for our experiments are available at https://github.com/ultramarineX/DashFusion.

cs.CV↗

HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios

Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal content separately, losing essential acoustic information that degrades output quality while requiring significant computational resources. To overcome these limitations, we propose HQ-SVC, an efficient framework for high-quality zero-shot SVC. HQ-SVC first extracts jointly content and speaker features using a decoupled codec. It then enhances fidelity through pitch and volume modeling, preserving critical acoustic information typically lost in separate modeling approaches, and progressively refines outputs via differentiable signal processing and diffusion techniques. Evaluations confirm HQ-SVC significantly outperforms state-of-the-art zero-shot SVC methods in conversion quality and efficiency. Beyond voice conversion, HQ-SVC achieves superior voice naturalness compared to specialized audio super-resolution methods while natively supporting voice super-resolution tasks.

cs.SD↗

$CP$ violation in two-body hadronic $Λ_b$ decays in the PQCD approach

We systematically investigate the $CP$-averaged branching ratios and $CP$ violations (CPVs) for the two-body hadronic decays $Λ_b\to ph$, where $h$ runs through the mesons $π^-$, $ρ^-$, $a_1^-(1260)$, $K^-$, $K^{\ast -}$, $K_1^-(1270)$ and $K_1^-(1400)$, in the perturbative QCD approach to order $α_s^2$ in the strong coupling. Various topological amplitudes are obtained by incorporating subleading-twist hadron distribution amplitudes, which exhibit reasonable hierarchical patterns, sizable strong phases, and non-negligible higher-power corrections. The predicted direct CPVs in $Λ_b\to pπ^-,pK^-$, different from those in similar $B$ meson decays, are as small as the current data. The low CPV in $Λ_b\to pπ^-$ results from the cancellation between the $S$- and $P$-wave CPVs, while the one in $Λ_b\to pK^-$ is determined by the tiny $S$-wave CPV. However, individual partial-wave CPVs can exceed $10\%$, consistent with direct CPVs in $B$ meson decays. The CPVs in the $Λ_b\to pK_1^-(1270),pK_1^-(1400)$ channels are relatively larger. In particular, CPVs above $20\%$ appear in the up-down asymmetries associated with the final-state angular distributions of $Λ_b\to pK_1^-(1270),pK_1^-(1400)$, followed by the secondary $K_1\to Kππ$ decays. These observables offer promising prospects for firmly establishing baryon CPVs. The decay asymmetry parameters of $Λ_b\to ph$ are also predicted for future experimental confrontations.

hep-ph↗

Efficient and Robust Spatial-to-Fiber Coupling forMultimode Quantum Networks via CascadedAdaptive Feedback Control

Duan-Lukin-Cirac-Zoller (DLCZ)-based multimodequantum networks rely on efficient spatial-to-fiber coupling, yetenvironmental perturbations compromise this performance. Wedevelop a cascaded adaptive feedback control system integratedinto the quantum entanglement source preparation path.Leveraging a power-feedback hillclimbing algorithm, itdynamically regulates piezoelectric-actuated mirrors to achieveautonomous multi-dimensional beam alignment, Experimentsshow it rapidly boosts single-mode fiber (SMF) coupling efficieneyto over 70% within 20 seconds and entering the most efficient andstable transmission state after 75 seconds.Importantly, it enhancesthe stability of the atom-photon interfacecritical for quantumlight-matter interactionsproviding a practical framework forefficient, robust spatial light transmission in scalable quantumnetworks.

quant-ph↗

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from incomplete speech, inaccurate or missing timestamps, and limited real-world relevance. To address these problems, we propose an automated framework for generating large-scale paralinguistic data and apply it to construct the SynParaSpeech dataset. The dataset comprises 6 paralinguistic categories with 118.75 hours of data and precise timestamps, all derived from natural conversational speech. Our contributions lie in introducing the first automated method for constructing large-scale paralinguistic datasets and releasing the SynParaSpeech corpus, which advances speech generation through more natural paralinguistic synthesis and enhances speech understanding by improving paralinguistic event detection. The dataset and audio samples are available at https://github.com/ShawnPi233/SynParaSpeech.

eess.AS↗

Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios. To advance CMs for real-world deepfake detection, we first propose the Fake Speech Wild (FSW) dataset, which includes 254 hours of real and deepfake audio from four different media platforms, focusing on social media. As CMs, we establish a benchmark using public datasets and advanced selfsupervised learning (SSL)-based CMs to evaluate current CMs in real-world scenarios. We also assess the effectiveness of data augmentation strategies in enhancing CM robustness for detecting deepfake speech on social media. Finally, by augmenting public datasets and incorporating the FSW training set, we significantly advanced real-world deepfake audio detection performance, achieving an average equal error rate (EER) of 3.54% across all evaluation sets.

cs.SD↗

Video Demoireing using Focused-Defocused Dual-Camera System

Moire patterns, unwanted color artifacts in images and videos, arise from the interference between spatially high-frequency scene contents and the spatial discrete sampling of digital cameras. Existing demoireing methods primarily rely on single-camera image/video processing, which faces two critical challenges: 1) distinguishing moire patterns from visually similar real textures, and 2) preserving tonal consistency and temporal coherence while removing moire artifacts. To address these issues, we propose a dual-camera framework that captures synchronized videos of the same scene: one in focus (retaining high-quality textures but may exhibit moire patterns) and one defocused (with significantly reduced moire patterns but blurred textures). We use the defocused video to help distinguish moire patterns from real texture, so as to guide the demoireing of the focused video. We propose a frame-wise demoireing pipeline, which begins with an optical flow based alignment step to address any discrepancies in displacement and occlusion between the focused and defocused frames. Then, we leverage the aligned defocused frame to guide the demoireing of the focused frame using a multi-scale CNN and a multi-dimensional training loss. To maintain tonal and temporal consistency, our final step involves a joint bilateral filter to leverage the demoireing result from the CNN as the guide to filter the input focused frame to obtain the final output. Experimental results demonstrate that our proposed framework largely outperforms state-of-the-art image and video demoireing methods.

cs.CV↗

Deep Learning Approaches for Multimodal Intent Recognition: A Survey

Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer interaction, the field has evolved through deep learning and multimodal approaches, incorporating data from audio, vision, and physiological signals. Recently, the introduction of Transformer-based models has led to notable breakthroughs in this domain. This article surveys deep learning methods for intent recognition, covering the shift from unimodal to multimodal techniques, relevant datasets, methodologies, applications, and current challenges. It provides researchers with insights into the latest developments in multimodal intent recognition (MIR) and directions for future research.

cs.CL↗

ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation

The Segment Anything Model (SAM), with its prompt-driven paradigm, exhibits strong generalization in generic segmentation tasks. However, applying SAM to remote sensing (RS) images still faces two major challenges. First, manually constructing precise prompts for each image (e.g., points or boxes) is labor-intensive and inefficient, especially in RS scenarios with dense small objects or spatially fragmented distributions. Second, SAM lacks domain adaptability, as it is pre-trained primarily on natural images and struggles to capture RS-specific semantics and spatial characteristics, especially when segmenting novel or unseen classes. To address these issues, inspired by few-shot learning, we propose ViRefSAM, a novel framework that guides SAM utilizing only a few annotated reference images that contain class-specific objects. Without requiring manual prompts, ViRefSAM enables automatic segmentation of class-consistent objects across RS images. Specifically, ViRefSAM introduces two key components while keeping SAM's original architecture intact: (1) a Visual Contextual Prompt Encoder that extracts class-specific semantic clues from reference images and generates object-aware prompts via contextual interaction with target images; and (2) a Dynamic Target Alignment Adapter, integrated into SAM's image encoder, which mitigates the domain gap by injecting class-specific semantics into target image features, enabling SAM to dynamically focus on task-relevant regions. Extensive experiments on three few-shot segmentation benchmarks, including iSAID-5$^i$, LoveDA-2$^i$, and COCO-20$^i$, demonstrate that ViRefSAM enables accurate and automatic segmentation of unseen classes by leveraging only a few reference images and consistently outperforms existing few-shot segmentation methods across diverse datasets.

cs.CV↗

Semileptonic baryon decays $Ξ_b\rightarrow Ξ_c \ell^- \barν_\ell $ in perturbative QCD

We perform a detailed analysis of the semileptonic $Ξ_b\rightarrow Ξ_c \ell^- \barν_\ell$ decays within the perturbative QCD framework. In our study, the $Ξ_b\rightarrow Ξ_c$ transition form factors are calculated using several popular models for baryonic light-cone distribution amplitudes. These form factors are then employed to analyze a range of observable quantities for the semileptonic processes via the helicity formalism. Our work presents predictions for the branching fractions of these decays for both the $τ$ and $e$ channels. Notably, the obtained lepton flavor universality ratio, $\mathcal{R}_{Ξ_c}\approx 0.3$, may offer new insights into the $\mathcal{R}^{(*)}$ puzzle. Furthermore, we investigate various angular observables, such as forward-backward asymmetries, lepton-side convexity parameters, and polarization asymmetries, which provide complementary information regarding potential new physics in $b$-baryonic semileptonic transitions. The numerical results for these angular observables are presented as both functions of $q^2$ and as averaged values. We observe that the lepton mass plays a significant role in shaping the angular distributions, affecting most of the observables under consideration. These results are expected to be valuable for both current and future experimental investigations of semileptonic heavy-to-heavy baryon decays.

hep-ph↗

Establishing CP Violation in $b$-Baryon Decays

It is a long-standing puzzle why the {\it CP} violation (CPV) in the baryon system has not yet been definitively established as in the meson one. We demonstrate that individual partial-wave CPV in the $Λ_b\to pπ^-$ and $pK^-$ decays can exceed $10\%$, but the destruction between the partial waves (the suppression by the small partial-wave weight) results in a small net direct CPV in the former (the latter) as measured currently. Our finding highlights the different dynamics responsible for CPVs in baryon and meson decays. We propose to probe the CPV observables associated with the angular distributions of the $Λ_b\to pa_1(1260)$, $pK_1(1270)$ decay products, which are large enough for being identified experimentally.

hep-ph↗

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT.

cs.HC↗