Searcharxiv⌕ Search

arXiv subjects

Dong Yu

Publications and source records attributed to Dong Yu.

410 records · Page 23Linked to original sources

Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training

Although great progresses have been made in automatic speech recognition (ASR), significant performance degradation is still observed when recognizing multi-talker mixed speech. In this paper, we propose and evaluate several architectures to address this problem under the assumption that only a single channel of mixed signal is available. Our technique extends permutation invariant training (PIT) by introducing the front-end feature separation module with the minimum mean square error (MSE) criterion and the back-end recognition module with the minimum cross entropy (CE) criterion. More specifically, during training we compute the average MSE or CE over the whole utterance for each possible utterance-level output-target assignment, pick the one with the minimum MSE or CE, and optimize for that assignment. This strategy elegantly solves the label permutation problem observed in the deep learning based multi-talker mixed speech separation and recognition systems. The proposed architectures are evaluated and compared on an artificially mixed AMI dataset with both two- and three-talker mixed speech. The experimental results indicate that our proposed architectures can cut the word error rate (WER) by 45.0% and 25.0% relatively against the state-of-the-art single-talker speech recognition system across all speakers when their energies are comparable, for two- and three-talker mixed speech, respectively. To our knowledge, this is the first work on the multi-talker mixed speech recognition on the challenging speaker-independent spontaneous large vocabulary continuous speech task.

cs.SD↗

Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks

In this paper we propose the utterance-level Permutation Invariant Training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep learning based solution for speaker independent multi-talker speech separation. Specifically, uPIT extends the recently proposed Permutation Invariant Training (PIT) technique with an utterance-level cost function, hence eliminating the need for solving an additional permutation problem during inference, which is otherwise required by frame-level PIT. We achieve this using Recurrent Neural Networks (RNNs) that, during training, minimize the utterance-level separation error, hence forcing separated frames belonging to the same speaker to be aligned to the same output stream. In practice, this allows RNNs, trained with uPIT, to separate multi-talker mixed speech without any prior knowledge of signal duration, number of speakers, speaker identity or gender. We evaluated uPIT on the WSJ0 and Danish two- and three-talker mixed-speech separation tasks and found that uPIT outperforms techniques based on Non-negative Matrix Factorization (NMF) and Computational Auditory Scene Analysis (CASA), and compares favorably with Deep Clustering (DPCL) and the Deep Attractor Network (DANet). Furthermore, we found that models trained with uPIT generalize well to unseen speakers and languages. Finally, we found that a single model, trained with uPIT, can handle both two-speaker, and three-speaker speech mixtures.

cs.SD↗

Recognizing Multi-talker Speech with Permutation Invariant Training

In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) for automatic speech recognition (ASR). In PIT-ASR, we compute the average cross entropy (CE) over all frames in the whole utterance for each possible output-target assignment, pick the one with the minimum CE, and optimize for that assignment. PIT-ASR forces all the frames of the same speaker to be aligned with the same output layer. This strategy elegantly solves the label permutation problem and speaker tracing problem in one shot. Our experiments on artificially mixed AMI data showed that the proposed approach is very promising.

cs.SD↗

Deep Embedding Forest: Forest-based Serving with Deep Embedding Features

Deep Neural Networks (DNN) have demonstrated superior ability to extract high level embedding vectors from low level features. Despite the success, the serving time is still the bottleneck due to expensive run-time computation of multiple layers of dense matrices. GPGPU, FPGA, or ASIC-based serving systems require additional hardware that are not in the mainstream design of most commercial applications. In contrast, tree or forest-based models are widely adopted because of low serving cost, but heavily depend on carefully engineered features. This work proposes a Deep Embedding Forest model that benefits from the best of both worlds. The model consists of a number of embedding layers and a forest/tree layer. The former maps high dimensional (hundreds of thousands to millions) and heterogeneous low-level features to the lower dimensional (thousands) vectors, and the latter ensures fast serving. Built on top of a representative DNN model called Deep Crossing, and two forest/tree-based models including XGBoost and LightGBM, a two-step Deep Embedding Forest algorithm is demonstrated to achieve on-par or slightly better performance as compared with the DNN counterpart, with only a fraction of serving time on conventional hardware. After comparing with a joint optimization algorithm called partial fuzzification, also proposed in this paper, it is concluded that the two-step Deep Embedding Forest has achieved near optimal performance. Experiments based on large scale data sets (up to 1 billion samples) from a major sponsored search engine proves the efficacy of the proposed model.

cs.LG↗

Permutation Invariant Training of Deep Models for Speaker-Independent Multi-talker Speech Separation

We propose a novel deep learning model, which supports permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from most of the prior arts that treat speech separation as a multi-class regression problem and the deep clustering technique that considers it a segmentation (or clustering) problem, our model optimizes for the separation regression error, ignoring the order of mixing sources. This strategy cleverly solves the long-lasting label permutation problem that has prevented progress on deep learning based techniques for speech separation. Experiments on the equal-energy mixing setup of a Danish corpus confirms the effectiveness of PIT. We believe improvements built upon PIT can eventually solve the cocktail-party problem and enable real-world adoption of, e.g., automatic meeting transcription and multi-party human-computer interaction, where overlapping speech is common.

cs.CL↗

Very Strong Superconducting Proximity Effects in PbS Semiconductor Nanowires

We report the fabrication of strongly coupled nanohybrid superconducting junctions using PbS semiconductor nanowires and Pb0.5In0.5 superconducting electrodes. The maximum supercurrent in the junction reaches up to ~15 μA at 0.3 K, which is the highest value ever observed in semiconductor-nanowire-based superconducting junctions. The observation of microwave-induced constant voltage steps confirms the existence of genuine Josephson coupling through the nanowire. Monotonic suppression of the critical current under an external magnetic field is also in good agreement with the narrow junction model. The temperature-dependent stochastic distribution of the switching current exhibits a crossover from phase diffusion to a thermal activation process as the temperature decreases. These strongly coupled nanohybrid superconducting junctions would be advantageous to the development of gate-tunable superconducting quantum information devices.

cond-mat.supr-con↗

Highway Long Short-Term Memory RNNs for Distant Speech Recognition

In this paper, we extend the deep long short-term memory (DLSTM) recurrent neural networks by introducing gated direct connections between memory cells in adjacent layers. These direct links, called highway connections, enable unimpeded information flow across different layers and thus alleviate the gradient vanishing problem when building deeper LSTMs. We further introduce the latency-controlled bidirectional LSTMs (BLSTMs) which can exploit the whole history while keeping the latency under control. Efficient algorithms are proposed to train these novel networks using both frame and sequence discriminative criteria. Experiments on the AMI distant speech recognition (DSR) task indicate that we can train deeper LSTMs and achieve better improvement from sequence training with highway LSTMs (HLSTMs). Our novel model obtains $43.9/47.7\%$ WER on AMI (SDM) dev and eval sets, outperforming all previous works. It beats the strong DNN and DLSTM baselines with $15.7\%$ and $5.3\%$ relative improvement respectively.

cs.NE↗

Spin Generation Via Bulk Spin Current in Three Dimensional Topological Insulators

To date, spin generation in three-dimensional topological insulators is primarily modeled as a single-surface phenomenon, attributed to the momentum-spin locking on each individual surface. In this article we propose a mechanism of spin generation where the role of the insulating yet topologically non-trivial bulk becomes explicit: an external electric field creates a transverse pure spin current through the bulk of a three-dimensional topological insulator, which transports spins between the top and bottom surfaces. Under sufficiently high surface disorder, the spin relaxation time can be extended via the Dyakonov-Perel mechanism. Consequently both the spin generation efficiency and surface conductivity are largely enhanced. Numerical simulation confirms that this spin generation mechanism originates from the unique topological connection of the top and bottom surfaces and is absent in other two dimensional systems such as graphene, even though they possess a similar Dirac cone-type dispersion.

cond-mat.mtrl-sci↗

Gate-tunable superconducting quantum interference devices of PbS nanowires

We report on the fabrication and electrical transport properties of gate-tunable superconducting quantum interference devices (SQUIDs), made of semiconducting PbS nanowire contacted with PbIn superconducting electrodes. Applied with a magnetic field perpendicular to the plane of the nano-hybrid SQUID, periodic oscillations of the critical current due to the flux quantization in SQUID are observed up to T = 4.0 K. Nonsinusoidal current-phase relationship is obtained as a function of temperature and gate voltage, which is consistent with a short and diffusive junction model.

cond-mat.mes-hall↗

Hot Carrier Trapping Induced Negative Photoconductance in InAs Nanowires toward Novel Nonvolatile Memory

We report a novel negative photoconductivity (NPC) mechanism in n-type indium arsenide nanowires (NWs). Photoexcitation significantly suppresses the conductivity with a gain up to 10^5. The origin of NPC is attributed to the depletion of conduction channels by light assisted hot electron trapping, supported by gate voltage threshold shift and wavelength dependent photoconductance measurements. Scanning photocurrent microscopy excludes the possibility that NPC originates from the NW/metal contacts and reveals a competing positive photoconductivity. The conductivity recovery after illumination substantially slows down at low temperature, indicating a thermally activated detrapping mechanism. At 78 K, the spontaneous recovery of the conductance is completely quenched, resulting in a reversible memory device which can be switched by light and gate voltage pulses. The novel NPC based optoelectronics may find exciting applications in photodetection and nonvolatile memory with low power consumption.

cond-mat.mes-hall↗

Prediction-Adaptation-Correction Recurrent Neural Networks for Low-Resource Language Speech Recognition

In this paper, we investigate the use of prediction-adaptation-correction recurrent neural networks (PAC-RNNs) for low-resource speech recognition. A PAC-RNN is comprised of a pair of neural networks in which a {\it correction} network uses auxiliary information given by a {\it prediction} network to help estimate the state probability. The information from the correction network is also used by the prediction network in a recurrent loop. Our model outperforms other state-of-the-art neural networks (DNNs, LSTMs) on IARPA-Babel tasks. Moreover, transfer learning from a language that is similar to the target language can help improve performance further.

cs.CL↗

Broadband Quantum Efficiency Enhancement in High Index Nanowires Resonators

Light trapping in sub-wavelength semiconductor nanowires (NWs) offers a promising approach to simultaneously reducing material consumption and enhancing photovoltaic performance. Nevertheless, the absorption efficiency of a NW, defined by the ratio of optical absorption cross section to the NW diameter, lingers around 1 in existing NW photonic devices, and the absorption enhancement suffers from a narrow spectral width. Here, we show that the absorption efficiency can be significantly improved in NWs with higher refractive indices, by an experimental observation of up to 350% external quantum efficiency (EQE) in lead sulfide (PbS) NW resonators, a 3-fold increase compared to Si NWs. Furthermore, broadband absorption enhancement is achieved in single tapered NWs, where light of various wavelengths is absorbed at segments with different diameters analogous to a tandem solar cell. Overall, the single NW Schottky junction solar cells benefit from optical resonance, near bandgap open circuit voltage, and long minority carrier diffusion length, demonstrating power conversion efficiency (PCE) comparable to single Si NW coaxial p-n junction cells11, but with much simpler fabrication processes.

cond-mat.mes-hall↗

Feature Learning in Deep Neural Networks - Studies on Speech Recognition Tasks

Recent studies have shown that deep neural networks (DNNs) perform significantly better than shallow networks and Gaussian mixture models (GMMs) on large vocabulary speech recognition tasks. In this paper, we argue that the improved accuracy achieved by the DNNs is the result of their ability to extract discriminative internal representations that are robust to the many sources of variability in speech signals. We show that these representations become increasingly insensitive to small perturbations in the input with increasing network depth, which leads to better speech recognition performance with deeper networks. We also show that DNNs cannot extrapolate to test samples that are substantially different from the training examples. If the training data are sufficiently representative, however, internal features learned by the DNN are relatively stable with respect to speaker differences, bandwidth differences, and environment distortion. This enables DNN-based recognizers to perform as well or better than state-of-the-art systems based on GMMs or shallow networks without the need for explicit model adaptation or feature normalization.

cs.LG↗

Variable range hopping conduction in semiconductor nanocrystal solids

The temperature and electrical field dependent conductivity of n-type CdSe nanocrystal thin films is investigated. In the low electrical field regime, the conductivity follows G ~ exp(-(T*/T)^0.5) in the temperature range 10K<T<120K. At high electrical field, the conductivity is strongly field dependent. At 4K, the conductance increases by eight orders of magnitude over one decade of bias. At very high field, conductivity is temperature-independent with G ~ exp(-(E*/E)^0.5). The complete behavior is very well described by variable range hopping with Coulomb gap.

cond-mat.mtrl-sci↗