SearcharxivSearch

arXiv subjects

Shengqiang Liu

Publications and source records attributed to Shengqiang Liu.

7 recordsLinked to original sources

Turbo your multi-modal classification with contrastive learning

Contrastive learning has become one of the most impressive approaches for multi-modal representation learning. However, previous multi-modal works mainly focused on cross-modal understanding, ignoring in-modal contrastive learning, which limits the representation of each modality. In this paper, we propose a novel contrastive learning strategy, called $Turbo$, to promote multi-modal understanding by joint in-modal and cross-modal contrastive learning. Specifically, multi-modal data pairs are sent through the forward pass twice with different hidden dropout masks to get two different representations for each modality. With these representations, we obtain multiple in-modal and cross-modal contrastive objectives for training. Finally, we combine the self-supervised Turbo with the supervised multi-modal classification and demonstrate its effectiveness on two audio-text classification tasks, where the state-of-the-art performance is achieved on a speech emotion recognition benchmark dataset.

cs.LG

M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection

With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with complex sound sources, the voice assistant must classify utterances as device-oriented or non-device-oriented. The dual-encoder structure, which is jointly modeled by text and speech, has become the paradigm of device-directed speech detection. However, in practice, these models often produce incorrect predictions for unaligned input pairs due to the unavoidable errors of automatic speech recognition (ASR).To address this challenge, we propose M$^{3}$V, a multi-modal multi-view approach for device-directed speech detection, which frames we frame the problem as a multi-view learning task that introduces unimodal views and a text-audio alignment view in the network besides the multi-modal. Experimental results show that M$^{3}$V significantly outperforms models trained using only single or multi-modality and surpasses human judgment performance on ASR error data for the first time.

cs.SD

DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training

Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on various downstream tasks of IVAs with pre-trained audio models and text models. However, these models are pre-trained independently and usually on tasks different from target domains, resulting in sub-optimal modality representations for downstream tasks. Moreover, in many domains, collecting enough language-audio pairs is extremely hard, and transcribing raw audio also requires high professional skills, making it difficult or even infeasible to joint pre-training. To address these painpoints, we propose DSCLAP, a simple and effective framework that enables language-audio pre-training with only raw audio signal input. Specifically, DSCLAP converts raw audio signals into text via an ASR system and combines a contrastive learning objective and a language-audio matching objective to align the audio and ASR transcriptions. We pre-train DSCLAP on 12,107 hours of in-vehicle domain audio. Empirical results on two downstream tasks show that while conceptually simple, DSCLAP significantly outperforms the baseline models in all metrics, showing great promise for domain-specific IVAs applications.

cs.SD

Wave propagation for a discrete diffusive vaccination epidemic model with bilinear incidence

The aim of the current paper is to study the existence of traveling wave solutions (TWS) for a vaccination epidemic model with bilinear incidence. The existence result is determined by the basic reproduction number $\Re_0$. More specifically, the system admits a nontrivial TWS when $\Re_0>1$ and $c \geq c^*$, where $c^*$ is the critical wave speed. We also found that the TWS is connecting two different equilibria by constructing Lyapunov functional. Lastly, we give some biological explanations from the perspective of epidemiology.

math.DS

Global Dynamics of a Predator-Prey Model with State-Dependent Maturation-Delay

In this paper, a stage structured predator-prey model with general nonlinear type of functional response is established and analyzed. The state-dependent time delay (hereafter SDTD) is the time taken from predator's birth to its maturity, formatted as a monotonical (ly) increasing, continuous(ly) differentiable and bounded function on the number of mature predator. The model is quite different from many previous models with SDTD, in the sense that the derivative of delay on the time is involved in the model. First, we have shown that for a large class of commonly used types of functional responses, including Holling types I, II and III, Beddington-DeAngelis-type (hereafter BD-type), etc, the predator coexists with the prey permanently if and only if the predator's net reproduction number is larger than one unit; Secondly, we have discussed the local stability of the equilibria of the model; Finally, for the special case of BD-type functional response, we claim that if the system is permanent, that is, the derivative of SDTD on the state is small enough and the predator interference is large enough, then the coexistence equilibrium is globally asymptotically stable.

math.DS

Traveling wave solutions for a class of discrete diffusive SIR epidemic model

This paper is concerned with the conditions of existence and nonexistence of traveling wave solutions (TWS) for a class of discrete diffusive epidemic models. We find that the existence of TWS is determined by the so-called basic reproduction number and the critical wave speed: When the basic reproduction number R0 greater than 1, there exists a critical wave speed c* > 0, such that for each c >= c * the system admits a nontrivial TWS and for c < c* there exists no nontrivial TWS for the system. In addition, the boundary asymptotic behaviour of TWS is obtained by constructing a suitable Lyapunov functional and employing Lebesgue dominated convergence theorem. Finally, we apply our results to two discrete diffusive epidemic models to verify the existence and nonexistence of TWS.

math.DS

Threshold dynamics and ergodicity of an SIRS epidemic model with Markovian switching

This paper studies the spread dynamics of a stochastic SIRS epidemic model with nonlinear incidence and varying population size, which is formulated as a piecewise deterministic Markov process. A threshold dynamic determined by the basic reproduction number $\mathcal{R}_{0}$ is established: the disease can be eradicated almost surely if $\mathcal{R}_{0}<1$, while the disease persists almost surely if $\mathcal{R}_{0}>1$. The existing method for analyzing ergodic behavior of population systems has been generalized. The modified method weakens the required conditions and has no limitations for both the number of environmental regimes and the dimension of the considered system. When $\mathcal{R}_{0}>1$, the existence of a stationary probability measure is obtained. Furthermore, with the modified method, the global attractivity of the $Ω$-limit set of the system and the convergence in total variation to the stationary measure are both demonstrated under a mild extra condition.

math.DS