SearcharxivSearch

arXiv subjects

Jingjing Yin

Publications and source records attributed to Jingjing Yin.

9 recordsLinked to original sources

On the Frobenius Number and Genus of a Collection of Semigroups Generalizing Repunit Numerical Semigroups

Let $A=(a_1, a_2, \ldots, a_n)$ be a sequence of relative prime positive integers with $a_i\geq 2$. The Frobenius number $F(A)$ is the largest integer not belonging to the numerical semigroup $\langle A\rangle$ generated by $A$. The genus $g(A)$ is the number of positive integer elements not in $\langle A\rangle$. The Frobenius problem is to determine $F(A)$ and $g(A)$ for a given sequence $A$. In this paper, we study the Frobenius problem of $A=\left(a,h_1a+b_1d,h_2a+b_2d,\ldots,h_ka+b_kd\right)$ with some restrictions. An innovation is that $d$ can be a negative integer. In particular, when $A=\left(a,ba+d,b^2a+\frac{b^2-1}{b-1}d,\ldots,b^ka+\frac{b^k-1}{b-1}d\right)$, we obtain formulas for $F(A)$ and $g(A)$ when $a\geq k-1-\frac{d-1}{b-1}$. Our formulas simplify further for some special cases, such as Mersenne, Thabit, and repunit numerical semigroups. Finally, we partially solve an open problem for the Proth numerical semigroup.

math.NT

Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin AudioLLM, a series of techniques and models, mainly including Takin TTS, Takin VC, and Takin Morphing, specifically designed for audiobook production. These models are capable of zero-shot speech production, generating high-quality speech that is nearly indistinguishable from real human speech and facilitating individuals to customize the speech content according to their own needs. Specifically, we first introduce Takin TTS, a neural codec language model that builds upon an enhanced neural speech codec and a multi-task training framework, capable of generating high-fidelity natural speech in a zero-shot way. For Takin VC, we advocate an effective content and timbre joint modeling approach to improve the speaker similarity, while advocating for a conditional flow matching based decoder to further enhance its naturalness and expressiveness. Last, we propose the Takin Morphing system with highly decoupled and advanced timbre and prosody modeling approaches, which enables individuals to customize speech production with their preferred timbre and prosody in a precise and controllable manner. Extensive experiments validate the effectiveness and robustness of our Takin AudioLLM series models. For detailed demos, please refer to https://everest-ai.github.io/takinaudiollm/.

cs.SD

MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition

Despite notable progress, speech emotion recognition (SER) remains challenging due to the intricate and ambiguous nature of speech emotion, particularly in wild world. While current studies primarily focus on recognition and generalization abilities, our research pioneers an investigation into the reliability of SER methods in the presence of semantic data shifts and explores how to exert fine-grained control over various attributes inherent in speech signals to enhance speech emotion modeling. In this paper, we first introduce MSAC-SERNet, a novel unified SER framework capable of simultaneously handling both single-corpus and cross-corpus SER. Specifically, concentrating exclusively on the speech emotion attribute, a novel CNN-based SER model is presented to extract discriminative emotional representations, guided by additive margin softmax loss. Considering information overlap between various speech attributes, we propose a novel learning paradigm based on correlations of different speech attributes, termed Multiple Speech Attribute Control (MSAC), which empowers the proposed SER model to simultaneously capture fine-grained emotion-related features while mitigating the negative impact of emotion-agnostic representations. Furthermore, we make a first attempt to examine the reliability of the MSAC-SERNet framework using out-of-distribution detection methods. Experiments on both single-corpus and cross-corpus SER scenarios indicate that MSAC-SERNet not only consistently outperforms the baseline in all aspects, but achieves superior performance compared to state-of-the-art SER approaches.

cs.SD

PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations in style diversity or falls short in terms of the intuitive and interpretability of style representation. In this study, we propose PromptVC, a novel style voice conversion approach that employs a latent diffusion model to generate a style vector driven by natural language prompts. Specifically, the style vector is extracted by a style encoder during training, and then the latent diffusion model is trained independently to sample the style vector from noise, with this process being conditioned on natural language prompts. To improve style expressiveness, we leverage HuBERT to extract discrete tokens and replace them with the K-Means center embedding to serve as the linguistic content, which minimizes residual style information. Additionally, we deduplicate the same discrete token and employ a differentiable duration predictor to re-predict the duration of each token, which can adapt the duration of the same linguistic content to different styles. The subjective and objective evaluation results demonstrate the effectiveness of our proposed system.

eess.AS

PP-MeT: a Real-world Personalized Prompt based Meeting Transcription System

Speaker-attributed automatic speech recognition (SA-ASR) improves the accuracy and applicability of multi-speaker ASR systems in real-world scenarios by assigning speaker labels to transcribed texts. However, SA-ASR poses unique challenges due to factors such as speaker overlap, speaker variability, background noise, and reverberation. In this study, we propose PP-MeT system, a real-world personalized prompt based meeting transcription system, which consists of a clustering system, target-speaker voice activity detection (TS-VAD), and TS-ASR. Specifically, we utilize target-speaker embedding as a prompt in TS-VAD and TS-ASR modules in our proposed system. In constrast with previous system, we fully leverage pre-trained models for system initialization, thereby bestowing our approach with heightened generalizability and precision. Experiments on M2MeT2.0 Challenge dataset show that our system achieves a cp-CER of 11.27% on the test set, ranking first in both fixed and open training conditions.

eess.AS

A Note on Generalized Repunit Numerical Semigroups

Let $A=(a_1, a_2, ..., a_n)$ be relative prime positive integers with $a_i\geq 2$. The Frobenius number $F(A)$ is the largest integer not belonging to the numerical semigroup $\langle A\rangle$ generated by $A$. The genus $g(A)$ is the number of positive integer elements that are not in $\langle A\rangle$. The Frobenius problem is to find $F(A)$ and $g(A)$ for a given sequence $A$. In this note, we study the Frobenius problem of $A=\left(a,ba+d,b^2a+\frac{b^2-1}{b-1}d,...,b^ka+\frac{b^k-1}{b-1}d\right)$ and obtain formulas for $F(A)$ and $g(A)$ when $a\geq k-1$. Our formulas simplifies further for some special cases, such as repunit, Mersenne and Thabit numerical semigroups. The idea is similar to that in [\cite{LiuXin23},arXiv:2306.03459].

math.NT

The Frobenius Formula for $A=(a,ha+d,ha+b_2d,...,ha+b_kd)$

Given relative prime positive integers $A=(a_1, a_2, ..., a_n)$, the Frobenius number $g(A)$ is the largest integer not representable as a linear combination of the $a_i$'s with nonnegative integer coefficients. We find the ``Stable" property introduced for the square sequence $A=(a,a+1,a+2^2,\dots, a+k^2)$ naturally extends for $A(a)=(a,ha+dB)=(a,ha+d,ha+b_2d,...,ha+b_kd)$. This gives a parallel characterization of $g(A(a))$ as a ``congruence class function" modulo $b_k$ when $a$ is large enough. For orderly sequence $B=(1,b_2,\dots,b_k)$, we find good bound for $a$. In particular we calculate $g(a,ha+dB)$ for $B=(1,2,b,b+1)$, $B=(1,2,b,b+1,2b)$, $B=(1,b,2b-1)$ and $B=(1,2,...,k,K)$. Our idea also applies to the case $B=(b_1,b_2,...,b_k)$, $b_1> 1$.

math.CO

HYBRIDFORMER: improving SqueezeFormer with hybrid attention and NSR mechanism

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attention (SA). In addition, limited by the large convolution kernel size, the local modeling ability of SqueezeFormer is insufficient. In this paper, we propose a novel method HybridFormer to improve SqueezeFormer in a fast and efficient way. Specifically, we first incorporate linear attention (LA) and propose a hybrid LASA paradigm to increase the model's inference speed. Second, a hybrid neural architecture search (NAS) guided structural re-parameterization (SRep) mechanism, termed NSR, is proposed to enhance the ability of the model to extract local interactions. Extensive experiments conducted on the LibriSpeech dataset demonstrate that our proposed HybridFormer can achieve a 9.1% relative word error rate (WER) reduction over SqueezeFormer on the test-other dataset. Furthermore, when input speech is 30s, the HybridFormer can improve the model's inference speed up to 18%. Our source code is available online.

eess.AS

LMEC: Learnable Multiplicative Absolute Position Embedding Based Conformer for Speech Recognition

This paper proposes a Learnable Multiplicative absolute position Embedding based Conformer (LMEC). It contains a kernelized linear attention (LA) module called LMLA to solve the time-consuming problem for long sequence speech recognition as well as an alternative to the FFN structure. First, the ELU function is adopted as the kernel function of our proposed LA module. Second, we propose a novel Learnable Multiplicative Absolute Position Embedding (LM-APE) based re-weighting mechanism that can reduce the well-known quadratic temporal-space complexity of softmax self-attention. Third, we use Gated Linear Units (GLU) to substitute the Feed Forward Network (FFN) for better performance. Extensive experiments have been conducted on the public LibriSpeech datasets. Compared to the Conformer model with cosFormer style linear attention, our proposed method can achieve up to 0.63% word-error-rate improvement on test-other and improve the inference speed by up to 13% (left product) and 33% (right product) on the LA module.

eess.AS