SearcharxivSearch

arXiv subjects

Ya Li

Publications and source records attributed to Ya Li.

At least 19 recordsLinked to original sources

Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering

Instruct-TTS systems expand structured style labels into natural-language training instructions through LLM rewriting, yet we find that over 40% of unconstrained rewrites contain semantic drift that corrupts supervision and weakens generalization. We formalize this problem as instruction supervision instability and propose a data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations. On the Chinese split of InstructTTSEval, our recipe raises instruction-following from 34.5% without fine-tuning and 51.0% with naive fine-tuning to 56.4%, while constrained rewriting reduces drift from 40.4% to 15.4%. Ablations confirm the three mechanisms are complementary, and the drift taxonomy may generalize to instruction-driven generation beyond TTS.

cs.SD

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fail to effectively extract and leverage task-relevant information from the incomplete multimodal data. To address this challenge, we propose a unified framework termed Mutual Information Disentanglement with uncertainty-Aware fuSion (MIDAS), which effectively restructures multimodal representations under incomplete conditions. MIDAS adopts a variational modeling strategy to represent each modality with multivariate Gaussian latent variables and further decomposes them into shared and exclusive factors. To obtain reliable representations, we design a minimax objective that minimizes the mutual information between shared and exclusive spaces for stable disentanglement, while maximizing the mutual information among shared spaces across modalities to enhance semantic alignment. In addition, an uncertainty-aware fusion mechanism is introduced, where posterior variance is leveraged as a reliability indicator to adaptively weight latent features during fusion, ensuring robust integration even when modalities are incomplete. Extensive experiments on three widely used datasets show that MIDAS achieves strong and consistent performance gains over competitive baselines across a wide range of incomplete settings, demonstrating its effectiveness and robustness for incomplete data scenarios.

cs.AI

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.

cs.SD

CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis

Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source-lyric-correlated cues; one source-following patch can propagate through AR history. We introduce CLASVS. Its State-Control-Transition (SCT) routing keeps target-lyric and reference-melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State-Control Grounding (PSCG) learns this contract through paired-edit-free, content-consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete-AR Vevo2 and reduces macro-PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous-AR operating point for score-annotation-free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/clasvs-demo/.

cs.SD

Polarization fractions and helicity-dependent CP asymmetries in $B_{(s)} \to \rho\rho, \rho K^\ast$ and $K^\ast K^\ast$ decays

In this paper, we present a phenomenological analysis of $B_{(s)} \to \rho\rho, \rho K^\ast$ and $K^\ast K^\ast$ decays using state-of-the-art perturbative QCD (pQCD) calculations. Our study is primarily motivated by recent polarization measurements from the LHCb and Belle II collaborations, which have significantly improved the precision of polarization fractions and enabled the first full determination of polarization-dependent CP asymmetries. This work extends the comprehensive pQCD study of charmless two-body $B$ decays reported in our previous paper [Chin. Phys. C 46 (2022) 123103], with a particular focus on polarization observables, especially the CP asymmetries in each helicity state, which reflect distinct orbital angular momentum configurations between the two vector mesons. Our predictions for the branching ratios and longitudinal polarization fractions in the $B^0 \to K^{\ast 0} {\bar K}^{\ast 0}$ and $B^+ \to \rho^0 K^{\ast +}$ modes are in good agreement with the new experimental data. However, the calculated longitudinal polarization fraction for $B_s \to K^{\ast 0} {\bar K}^{\ast 0}$ is significantly larger than the LHCb measurement. Moreover, the predicted (helicity-dependent) CP asymmetries in $B^+ \to \rho^0 K^{\ast +}$ are about $30 \%$ smaller than the observed values. These discrepancies point to a rich interplay between different topological amplitudes, highlighting the need for further theoretical investigation to resolve the long-standing polarization puzzle in two-body $B$ decays into vector mesons.

hep-ph

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions. In household environments, however, safety depends not only on recognizing objects, but also on how actions change the physical scene over time. Existing embodied safety evaluations largely focus on static risk recognition, unsafe instruction refusal, or final-state task completion. As a result, process-level safety failures induced by spatial relations such as support, containment, and proximity remain insufficiently studied. To address this gap, we introduce SAFERELBENCH, a spatial-relation-aware safety benchmark with 507 executable evaluation samples, including 248 spatial-relation samples and 259 non-spatial control samples. Using SAFERELBENCH to evaluate seven open- and closed-source VLM-driven embodied agents, we find a substantial gap between task success and process-level safety compliance: models often complete the requested task while violating process-level safety constraints. Unlike prior benchmarks, SAFERELBENCH explicitly tests whether agents satisfy safety conditions before risk-prone actions, making spatial relations a core dimension in embodied safety assessment. More broadly, our results show that safe embodied intelligence requires not only stronger perception and planning, but also reliable reasoning about how object relations shape risk during interaction.

cs.RO

Capacity Bounds and High-SNR Characterization for MIMO-OWC Channels Under Average-Power Constraint

This paper investigates the capacity of multipleinput multiple-output (MIMO) optical wireless communication (OWC) channels under a total average-power constraint. Since different nonnegative input vectors can be mapped to the same image vector and thus induce the same output distribution, we formulate a nonnegative basis pursuit (NN-BP) problem to identify the minimum-l1-norm input vector for each image vector. Based on the NN-BP characterization, we derive an equivalent expression for the channel capacity in terms of the image-vector distribution. We then establish computable lower and upper capacity bounds for both nT >= nR and nT < nR cases, and prove that the proposed bounds are asymptotically tight in the high signal-to-noise ratio (SNR) regime. Numerical results for indoor and outdoor OWC scenarios demonstrate that the proposed bounds improve upon existing ones and close the constant gap in the high-SNR regime.

cs.IT

AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ

Neural speech codecs have become the discrete interface between raw audio and speech language models, yet they remain optimized primarily for acoustic reconstruction fidelity, which leaves emotion-relevant cues vulnerable to being discarded during quantization, limiting the affective capacity of downstream models. We trace this degradation to two mechanisms: reconstruction-driven bit allocation under limited bitrate and cross-stream leakage in concatenation-based codecs, where acoustic gradients can overwrite nominally emotion-reserved dimensions. We propose AffectCodec, an emotion-preserving neural speech codec built on Block-Diagonal Residual Finite Scalar Quantization (BD-RFSQ). By imposing block-diagonal input and output projections over emotion and acoustic subspaces, BD-RFSQ transforms bit allocation from implicit and loss-driven to explicit and structurally guaranteed, while still preserving a flat token interface for downstream speech language models. AffectCodec further combines this structurally constrained quantizer with multi-granularity emotion conditioning and multi-rate training, enabling robust affect preservation at low bitrates. Experiments across multiple emotional speech benchmarks show that AffectCodec substantially improves emotion preservation, especially in the low-bitrate regime, while maintaining competitive acoustic quality and intelligibility. These results suggest that structurally protected quantization is an effective principle for preserving emotion-relevant information and may provide a general route toward attribute-aware neural speech compression.

cs.SD

Large CP violation in $\Lambda_b\rightarrow \Lambda D$ decays and extraction of the Cabibbo-Kobayashi-Maskawa angle $\gamma$

Motivated by the first observation of CP violation in $b$-baryon decays, the search for baryonic decays exhibiting large CP violation will be a primary focus in the coming years. We propose that significant CP-violating effects exist in the decay $\Lambda_b \to \Lambda D$, where $D$ denotes a CP eigenstate of the $D^0 - \bar{D}^0$ system. The predicted CP asymmetries for both the CP-even and CP-odd modes can reach magnitudes as large as $50\%$, making these decays promising targets for measurement at the LHCb experiment. Additionally, we predict for the first time several nonzero CP-violating observables associated with angular distribution parameters, providing valuable complementary information in the search for CP violation in baryon decays. Furthermore, we propose a novel strategy to extract the CKM angle $\gamma$ by combining data on angular distribution parameters and decay rates from the relevant channels. We emphasize that $\Lambda_b \rightarrow \Lambda D$ decays are among the most promising candidates for determining $\gamma$ in the baryon sector. Our findings may offer new insights for future theoretical and experimental investigations.

hep-ph

Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image

Reconstructing a complete 3D head from a single portrait remains challenging because existing methods still face a sharp quality-speed trade-off: high-fidelity pipelines often rely on multi-stage processing and per-subject optimization, while fast feed-forward models struggle with complete geometry and fine appearance details. To bridge this gap, we propose Any3DAvatar, a fast and high-quality method for single-image 3D Gaussian head avatar generation, whose fastest setting reconstructs a full head in under one second while preserving high-fidelity geometry and texture. First, we build AnyHead, a unified data suite that combines identity diversity, dense multi-view supervision, and realistic accessories, filling the main gaps of existing head data in coverage, full-head geometry, and complex appearance. Second, rather than sampling unstructured noise, we initialize from a Pl\"ucker-aware structured 3D Gaussian scaffold and perform one-step conditional denoising, formulating full-head reconstruction into a single forward pass while retaining high fidelity. Third, we introduce auxiliary view-conditioned appearance supervision on the same latent tokens alongside 3D Gaussian reconstruction, improving novel-view texture details at zero extra inference cost. Experiments show that Any3DAvatar outperforms prior single-image full-head reconstruction methods in rendering fidelity while remaining substantially faster.

cs.CV

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However, their effectiveness in low-resource languages remains fundamentally limited by the scarcity of transcribed speech. In practice, synthetic data has become the primary strategy for scaling SLMs in such settings, providing reliable phonetic supervision when real data is insufficient. In this work, we show that this reliance introduces a fundamental trade-off, which we term the Stability-Expressivity Gap: while synthetic data improves phonetic accuracy, it progressively suppresses prosodic variability, ultimately leading to a collapse of expressivity (Synthetic Erosion). To bridge this gap, we propose two self-alignment frameworks. Disentanglement-Guided Self-Alignment (DGSA) recovers expressivity for complex languages by exploiting prosody-timbre separation. For regimes where authentic references are exceptionally limited, Temperature-Driven Self-Critique (TDSC) stabilizes generation through automated exploration and filtering. Our approach outperforms strong commercial systems, including ElevenLabs and Gemini Pro, and enables the first zero-shot voice cloning capability for Lao.

cs.CL

VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents

With the growing demand for intelligent in-vehicle experiences, vehicle-based agents are evolving from simple assistants to long-term companions. This evolution requires agents to continuously model multi-user preferences and make reliable decisions in the face of inter-user preference conflicts and changing habits over time. However, existing benchmarks are largely limited to single-user, static question-answer settings, failing to capture the temporal evolution of preferences and the multi-user, tool-interactive nature of real vehicle environments. To address this gap, we introduce VehicleMemBench, a multi-user long-context memory benchmark built on an executable in-vehicle simulation environment. The benchmark evaluates tool use and memory by comparing the post-action environment state with a predefined target state, enabling objective and reproducible evaluation without LLM-based or human scoring. VehicleMemBench includes 23 tool modules, and each sample contains over 80 historical memory events. Experiments show that powerful models perform well on direct instruction tasks but struggle in scenarios involving memory evolution, particularly when user preferences change dynamically. Even advanced memory systems struggle to handle domain-specific memory requirements in this environment. These findings highlight the need for more robust and specialized memory management mechanisms to support long-term adaptive decision-making in real-world in-vehicle systems. To facilitate future research, we release the data and code.

cs.AI

Hello-Chat: Towards Realistic Social Audio Interactions

Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting in a robotic "read-speech" style that lacks the spontaneity and emotional resonance of real human interaction. In this report, we introduce Hello-Chat, an end-to-end audio language model designed for realistic social scenarios. By leveraging a massive dataset of real-life conversations and employing a modality-interleaved training strategy, Hello-Chat achieves a breakthrough in anthropomorphic generation. Experimental results show that our model not only reaches state-of-the-art (SOTA) performance on specific audio understanding tasks but also significantly outperforms existing baselines in prosodic naturalness and emotional alignment, paving the way for the next generation of empathetic AI agents.

cs.SD

Secrecy Capacity Analysis and Beamforming Optimization for MIMO-VLC Wiretap Channels

This paper investigates a multiple-input multipleoutput (MIMO) visible light communication (VLC) wiretap channel consisting of a transmitter, a legitimate receiver, and an eavesdropper. The optical input is subject to both peakand average-intensity constraints. By applying the generalized entropy-power inequality to truncated exponential inputs, we derive a novel closed-form expression for the achievable secrecy rate for general MIMO VLC configurations. To enhance transmission confidentiality, a fully-connected beamforming scheme is proposed, along with a low-complexity sub-connected alternative. Although the resulting beamforming design problems are nonconvex, they are efficiently addressed by transforming them into a sequence of convex subproblems solvable via the successive convex approximation framework. Numerical results demonstrate that the proposed schemes achieve significant secrecy performance improvements compared with the benchmark scheme.

cs.IT

$B_c$ meson decays into $S$-wave charmonium plus light meson pairs in the perturbative QCD approach

In this work, we explore the $P$-wave resonance contributions to the three-body charmonium decays of $B_c\to \Psi (V\to) P_1P_2$ using the perturbative QCD formalism at leading order, where $\Psi$ denotes a $S$-wave charmonium state, such as $\eta_c(1S,2S),J/\psi$, and $\psi(2S)$. Here, $P_1P_2$ represents a collinear $\pi\pi$ ($K\pi$) pair in the final state, which was primarily produced through the vector resonance $\rho(770)$ ($K^*(892)$ ). With the improved two-meson distribution amplitudes determined from our previous works, we examined the $CP$-averaged branching ratios and polarization fractions of the considered three-body decays. The longitudinal polarization fractions of the $B_c\to [J/\psi, \psi(2S)] (V\to) P_1P_2$ decays are found to be as large as $\sim 90\%$, since the transverse amplitudes from the dominant factorizable emission diagrams are always power suppressed with respect to the longitudinal ones. The direct $CP$ violations in $B_c\to \Psi (V\to) P_1P_2$ decays are predicted naturally to be zero as they solely receive contributions from tree diagrams. Several interesting relative ratios among the branching fractions of the concerned processes are investigated. In particular, the obtained ratio $R^{\rm PQCD}_{2\pi/\pi}\equiv \mathcal{B}(B^+_c \to J/\psi(\rho\to)\pi^+\pi^0)/{\mathcal{B}(B^+_c \to J/\psi\pi^+)}=2.67^{+0.21}_{-0.14}$ is consistent well with the LHCb measurement $R^{\rm exp}_{2\pi/\pi}=2.80\pm0.25$. Other similar ratios proposed in this work can be tested by LHCb experiments in the near future.

hep-ph

Semileptonic neutral current decays of $\Xi_b$ with dileptons or dineutrinos in the final state

We perform a detailed analysis of semileptonic $\Xi_b$ decays mediated by flavor-changing neutral currents ($b\to s$ and $b\to d$) with dilepton or dineutrino final states within the perturbative QCD framework. All independent form factors including vector, axial-vector, tensor, and pseudotensor currents are calculated and are used to analyze the decay branching fractions and angular distributions. Our numerical results for the branching fractions of $\Xi_b\to \Xi \ell^+\ell^-$ decays suggest they are within measurable reach for the LHCb experiment in the near future. Furthermore, we show that a measurement of the ratio $\mathcal{B}(\Xi_b^-\to \Sigma^- \mu^+\mu^-) / \mathcal{B}(\Xi_b^-\to \Xi^- \mu^+\mu^-)$ will allow for an independent determination of $|V_{td}/V_{ts}|$. For the case of unpolarized $\Xi_b$ baryons, we derive several angular observables, which can provide new and complementary constraints on Wilson coefficients in semileptonic FCNC transitions compared to those from mesonic decays. Finally, we present a combined analysis of dilepton and dineutrino channels, comparing various observables in detail. Our results offer further insights into the long-standing anomalies observed in $B$ meson decays.

hep-ph

The Renaissance of Expert Systems: Optical Recognition of Printed Chinese Jianpu Musical Scores with Lyrics

Large-scale optical music recognition (OMR) research has focused mainly on Western staff notation, leaving Chinese Jianpu (numbered notation) and its rich lyric resources underexplored. We present a modular expert-system pipeline that converts printed Jianpu scores with lyrics into machine-readable MusicXML and MIDI, without requiring massive annotated training data. Our approach adopts a top-down expert-system design, leveraging traditional computer-vision techniques (e.g., phrase correlation, skeleton analysis) to capitalize on prior knowledge, while integrating unsupervised deep-learning modules for image feature embeddings. This hybrid strategy strikes a balance between interpretability and accuracy. Evaluated on The Anthology of Chinese Folk Songs, our system massively digitizes (i) a melody-only collection of more than 5,000 songs (> 300,000 notes) and (ii) a curated subset with lyrics comprising over 1,400 songs (> 100,000 notes). The system achieves high-precision recognition on both melody (note-wise F1 = 0.951) and aligned lyrics (character-wise F1 = 0.931).

cs.CV

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment requires synchronizing both temporal and semantic information across modalities, while fusion involves integrating these aligned features into a unified representation. Existing methods often address alignment or fusion in isolation, leading to limitations in performance and efficiency. To tackle these issues, we propose a novel framework called Dual-stream Alignment with Hierarchical Bottleneck Fusion (DashFusion). Firstly, dual-stream alignment module synchronizes multimodal features through temporal and semantic alignment. Temporal alignment employs cross-modal attention to establish frame-level correspondences among multimodal sequences. Semantic alignment ensures consistency across the feature space through contrastive learning. Secondly, supervised contrastive learning leverages label information to refine the modality features. Finally, hierarchical bottleneck fusion progressively integrates multimodal information through compressed bottleneck tokens, which achieves a balance between performance and computational efficiency. We evaluate DashFusion on three datasets: CMU-MOSI, CMU-MOSEI, and CH-SIMS. Experimental results demonstrate that DashFusion achieves state-of-the-art performance across various metrics, and ablation studies confirm the effectiveness of our alignment and fusion techniques. The codes for our experiments are available at https://github.com/ultramarineX/DashFusion.

cs.CV