Searcharxiv⌕ Search

arXiv subjects

Gaofeng Cheng

Publications and source records attributed to Gaofeng Cheng.

21 records · Page 2Linked to original sources

A Layered Grouping Random Access Scheme Based on Dynamic Preamble Selection for Massive Machine Type Communications

Massive machine type communication (mMTC) has been identified as an important use case in Beyond 5G networks and future massive Internet of Things (IoT). However, for the massive multiple access in mMTC, there is a serious access preamble collision problem if the conventional 4-step random access (RA) scheme is employed. Consequently, a range of grantfree (GF) RA schemes were proposed. Nevertheless, if the number of cellular users (devices) significantly increases, both the energy and spectrum efficiency of the existing GF schemes still rapidly degrade owing to the much longer preambles required. In order to overcome this dilemma, a layered grouping strategy is proposed, where the cellular users are firstly divided into clusters based on their geographical locations, and then the users of the same cluster autonomously join in different groups by using optimum energy consumption (Opt-EC) based K-means algorithm. With this new layered cellular architecture, the RA process is divided into cluster load estimation phase and active group detection phase. Based on the state evolution theory of approximated message passing algorithm, a tight lower bound on the minimum preamble length for achieving a certain detection accuracy is derived. Benefiting from the cluster load estimation, a dynamic preamble selection (DPS) strategy is invoked in the second phase, resulting the required preambles with minimum length. As evidenced in our simulation results, this two-phase DPS aided RA strategy results in a significant performance improvement

cs.IT↗

Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture

Recently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contains the chunk self-attention encoder (chunk-SAE) and the monotonic truncated attention (MTA) based self-attention decoder (SAD). Firstly, the chunk-SAE splits the speech into isolated chunks. To reduce the computational cost and improve the performance, we propose the state reuse chunk-SAE. Sencondly, the MTA based SAD truncates the speech features monotonically and performs attention on the truncated features. To support the online recognition, we integrate the state reuse chunk-SAE and the MTA based SAD into online CTC/attention architecture. We evaluate the proposed online models on the HKUST Mandarin ASR benchmark and achieve a 23.66% character error rate (CER) with a 320 ms latency. Our online model yields as little as 0.19% absolute CER degradation compared with the offline baseline, and achieves significant improvement over our prior work on Long Short-Term Memory (LSTM) based online E2E models.

eess.AS↗

Utterance-level Permutation Invariant Training with Latency-controlled BLSTM for Single-channel Multi-talker Speech Separation

Utterance-level permutation invariant training (uPIT) has achieved promising progress on single-channel multi-talker speech separation task. Long short-term memory (LSTM) and bidirectional LSTM (BLSTM) are widely used as the separation networks of uPIT, i.e. uPIT-LSTM and uPIT-BLSTM. uPIT-LSTM has lower latency but worse performance, while uPIT-BLSTM has better performance but higher latency. In this paper, we propose using latency-controlled BLSTM (LC-BLSTM) during inference to fulfill low-latency and good-performance speech separation. To find a better training strategy for BLSTM-based separation network, chunk-level PIT (cPIT) and uPIT are compared. The experimental results show that uPIT outperforms cPIT when LC-BLSTM is used during inference. It is also found that the inter-chunk speaker tracing (ST) can further improve the separation performance of uPIT-LC-BLSTM. Evaluated on the WSJ0 two-talker mixed-speech separation task, the absolute gap of signal-to-distortion ratio (SDR) between uPIT-BLSTM and uPIT-LC-BLSTM is reduced to within 0.7 dB.

cs.SD↗