Searcharxiv⌕ Search

arXiv subjects

Yaping Zhang

Publications and source records attributed to Yaping Zhang.

24 records · Page 2Linked to original sources

Sequential Convolutional Recurrent Neural Networks for Fast Automatic Modulation Classification

A novel and efficient end-to-end learning model for automatic modulation classification is proposed for wireless spectrum monitoring applications, which automatically learns from the time domain in-phase and quadrature data without requiring the design of hand-crafted expert features. With the intuition of convolutional layers with pooling serving as the role of front-end feature distillation and dimensionality reduction, sequential convolutional recurrent neural networks are developed to take complementary advantage of parallel computing capability of convolutional neural networks and temporal sensitivity of recurrent neural networks. Experimental results demonstrate that the proposed architecture delivers overall superior performance in signal to noise ratio range above -10~dB, and achieves significantly improved classification accuracy from 80\% to 92.1\% at high signal to noise ratio range, while drastically reduces the average training and prediction time by approximately 74% and 67%, respectively. Response patterns learned by the proposed architecture are visualized to better understand the physics of the model. Furthermore, a comparative study is performed to investigate the impacts of various sequential convolutional recurrent neural network structure settings on classification performance. A representative sequential convolutional recurrent neural network architecture with the two-layer convolutional neural network and subsequent two-layer long short-term memory neural network is developed to suggest the option for fast automatic modulation classification.

eess.SP↗

Infer Implicit Contexts in Real-time Online-to-Offline Recommendation

Understanding users' context is essential for successful recommendations, especially for Online-to-Offline (O2O) recommendation, such as Yelp, Groupon, and Koubei. Different from traditional recommendation where individual preference is mostly static, O2O recommendation should be dynamic to capture variation of users' purposes across time and location. However, precisely inferring users' real-time contexts information, especially those implicit ones, is extremely difficult, and it is a central challenge for O2O recommendation. In this paper, we propose a new approach, called Mixture Attentional Constrained Denoise AutoEncoder (MACDAE), to infer implicit contexts and consequently, to improve the quality of real-time O2O recommendation. In MACDAE, we first leverage the interaction among users, items, and explicit contexts to infer users' implicit contexts, then combine the learned implicit-context representation into an end-to-end model to make the recommendation. MACDAE works quite well in the real system. We conducted both offline and online evaluations of the proposed approach. Experiments on several real-world datasets (Yelp, Dianping, and Koubei) show our approach could achieve significant improvements over state-of-the-arts. Furthermore, online A/B test suggests a 2.9% increase for click-through rate and 5.6% improvement for conversion rate in real-world traffic. Our model has been deployed in the product of "Guess You Like" recommendation in Koubei.

cs.IR↗

Deep Segment Attentive Embedding for Duration Robust Speaker Verification

LSTM-based speaker verification usually uses a fixed-length local segment randomly truncated from an utterance to learn the utterance-level speaker embedding, while using the average embedding of all segments of a test utterance to verify the speaker, which results in a critical mismatch between testing and training. This mismatch degrades the performance of speaker verification, especially when the durations of training and testing utterances are very different. To alleviate this issue, we propose the deep segment attentive embedding method to learn the unified speaker embeddings for utterances of variable duration. Each utterance is segmented by a sliding window and LSTM is used to extract the embedding of each segment. Instead of only using one local segment, we use the whole utterance to learn the utterance-level embedding by applying an attentive pooling to the embeddings of all segments. Moreover, the similarity loss of segment-level embeddings is introduced to guide the segment attention to focus on the segments with more speaker discriminations, and jointly optimized with the similarity loss of utterance-level embeddings. Systematic experiments on Tongdun and VoxCeleb show that the proposed method significantly improves robustness of duration variant and achieves the relative Equal Error Rate reduction of 50% and 11.54% , respectively.

eess.AS↗

Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of ASR. However, more complex pipelines, more computations and even higher hardware costs (microphone array) are additionally consumed for this kind of methods. In addition, speech enhancement would result in speech distortions and mismatches to training. In this paper, we propose an adversarial training method to directly boost noise robustness of acoustic model. Specifically, a jointly compositional scheme of generative adversarial net (GAN) and neural network-based acoustic model (AM) is used in the training phase. GAN is used to generate clean feature representations from noisy features by the guidance of a discriminator that tries to distinguish between the true clean signals and generated signals. The joint optimization of generator, discriminator and AM concentrates the strengths of both GAN and AM for speech recognition. Systematic experiments on CHiME-4 show that the proposed method significantly improves the noise robustness of AM and achieves the average relative error rate reduction of 23.38% and 11.54% on the development and test set, respectively.

cs.SD↗

Electron irradiation induced reduction of the permittivity in chalcogenide glass (As2S3) thin film

We investigate the effect of electron beam irradiation on the dielectric properties of As2S3 Chalcogenide glass. By means of low-loss Electron Energy Loss Spectroscopy, we derive the permittivity function, its dispersive relation, and calculate the refractive index and absorption coefficients under the constant permeability approximation. The measured and calculated results show, to the best of our knowledge, a heretofore unseen phenomenon: the reduction in the permittivity of <40%, and consequently a modification of the refractive index follows, reducing it by 20%, hence suggesting a significant change on the optical properties of the material. The plausible physical phenomena leading to these observations are discussed in terms of the homopolar and heteropolar bond dynamics under high energy absorption.

cond-mat.mtrl-sci↗

Generation of J0-Bessel-Gauss Beam by an heterogeneous refractive index map

In this paper, we present the theoretical studies of a refractive index map to implement a Gauss to J0-Bessel-Gauss convertor. We theoretically demonstrate the viability of such device by solving the inverse electromagnetic problem. The computed conversion efficiency is 90%. The theoretical results, obtained from the beam conversion efficiency, self-regeneration, and propagation through an opaque obstruction; demonstrate that a 2D graded index map of the refractive index can be used to transform a Gauss beam into a J0-Bessel-Gauss beam. To the best of our knowledge, this is the first demonstration of such beam transformation by means of a 2D index-mapping which is fully integrable in silicon photonics based planar lightwave circuits (PLC). The concept device is significant for the eventual development of a new array of technologies, such as micro optical tweezers, optical traps, beam reshaping and non-linear beam diode lasers.

physics.optics↗