Searcharxiv⌕ Search

arXiv subjects

Kun Han

Publications and source records attributed to Kun Han.

49 records · Page 3Linked to original sources

A Hybrid Task-Oriented Dialog System with Domain and Task Adaptive Pretraining

This paper describes our submission for the End-to-end Multi-domain Task Completion Dialog shared task at the 9th Dialog System Technology Challenge (DSTC-9). Participants in the shared task build an end-to-end task completion dialog system which is evaluated by human evaluation and a user simulator based automatic evaluation. Different from traditional pipelined approaches where modules are optimized individually and suffer from cascading failure, we propose an end-to-end dialog system that 1) uses Generative Pretraining 2 (GPT-2) as the backbone to jointly solve Natural Language Understanding, Dialog State Tracking, and Natural Language Generation tasks, 2) adopts Domain and Task Adaptive Pretraining to tailor GPT-2 to the dialog domain before finetuning, 3) utilizes heuristic pre/post-processing rules that greatly simplify the prediction tasks and improve generalizability, and 4) equips a fault tolerance module to correct errors and inappropriate responses. Our proposed method significantly outperforms baselines and ties for first place in the official evaluation. We make our source code publicly available.

cs.CL↗

Spatial Context-Aware Self-Attention Model For Multi-Organ Segmentation

Multi-organ segmentation is one of most successful applications of deep learning in medical image analysis. Deep convolutional neural nets (CNNs) have shown great promise in achieving clinically applicable image segmentation performance on CT or MRI images. State-of-the-art CNN segmentation models apply either 2D or 3D convolutions on input images, with pros and cons associated with each method: 2D convolution is fast, less memory-intensive but inadequate for extracting 3D contextual information from volumetric images, while the opposite is true for 3D convolution. To fit a 3D CNN model on CT or MRI images on commodity GPUs, one usually has to either downsample input images or use cropped local regions as inputs, which limits the utility of 3D models for multi-organ segmentation. In this work, we propose a new framework for combining 3D and 2D models, in which the segmentation is realized through high-resolution 2D convolutions, but guided by spatial contextual information extracted from a low-resolution 3D model. We implement a self-attention mechanism to control which 3D features should be used to guide 2D segmentation. Our model is light on memory usage but fully equipped to take 3D contextual information into account. Experiments on multiple organ segmentation datasets demonstrate that by taking advantage of both 2D and 3D models, our method is consistently outperforms existing 2D and 3D models in organ segmentation accuracy, while being able to directly take raw whole-volume image data as inputs.

eess.IV↗

Phase diagram and superconducting dome of infinite-layer $\mathrm{Nd_{1-x}Sr_{x}NiO_{2}}$ thin films

Infinite-layer Nd1-xSrxNiO2 thin films with Sr doping level x from 0.08 to 0.3 were synthesized and investigated. We found a superconducting dome to be between 0.12 and 0.235 which is accompanied by a weakly insulating behaviour in both underdoped and overdoped regimes. The dome is akin to that in the electron-doped 214-type and infinite-layer cuprate superconductors. For x higher than 0.18, the normal state Hall coefficient ($R_{H}$) changes the sign from negative to positive as the temperature decreases. The temperature of the sign changes monotonically decreases with decreasing x from the overdoped side and approaches the superconducting dome at the mid-point, suggesting a reconstruction of the Fermi surface as the dopant concentration changes across the center of the dome.

cond-mat.supr-con↗

Learning Alignment for Multimodal Emotion Recognition from Speech

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech recognition techniques to generate text from speech and then apply natural language processing to analyze the sentiment. Further, emotion recognition will be beneficial from using audio-textual multimodal information, it is not trivial to build a system to learn from multimodality. One can build models for two input sources separately and combine them in a decision level, but this method ignores the interaction between speech and text in the temporal domain. In this paper, we propose to use an attention mechanism to learn the alignment between speech frames and text words, aiming to produce more accurate multimodal feature representations. The aligned multimodal features are fed into a sequential model for emotion recognition. We evaluate the approach on the IEMOCAP dataset and the experimental results show the proposed approach achieves the state-of-the-art performance on the dataset.

cs.CL↗

Learning Syntactic and Dynamic Selective Encoding for Document Summarization

Text summarization aims to generate a headline or a short summary consisting of the major information of the source text. Recent studies employ the sequence-to-sequence framework to encode the input with a neural network and generate abstractive summary. However, most studies feed the encoder with the semantic word embedding but ignore the syntactic information of the text. Further, although previous studies proposed the selective gate to control the information flow from the encoder to the decoder, it is static during the decoding and cannot differentiate the information based on the decoder states. In this paper, we propose a novel neural architecture for document summarization. Our approach has the following contributions: first, we incorporate syntactic information such as constituency parsing trees into the encoding sequence to learn both the semantic and syntactic information from the document, resulting in more accurate summary; second, we propose a dynamic gate network to select the salient information based on the context of the decoder state, which is essential to document summarization. The proposed model has been evaluated on CNN/Daily Mail summarization datasets and the experimental results show that the proposed approach outperforms baseline approaches.

cs.CL↗

Adversarial Multi-Binary Neural Network for Multi-class Classification

Multi-class text classification is one of the key problems in machine learning and natural language processing. Emerging neural networks deal with the problem using a multi-output softmax layer and achieve substantial progress, but they do not explicitly learn the correlation among classes. In this paper, we use a multi-task framework to address multi-class classification, where a multi-class classifier and multiple binary classifiers are trained together. Moreover, we employ adversarial training to distinguish the class-specific features and the class-agnostic features. The model benefits from better feature representation. We conduct experiments on two large-scale multi-class text classification tasks and demonstrate that the proposed architecture outperforms baseline approaches.

cs.CL↗

Selective Attention Encoders by Syntactic Graph Convolutional Networks for Document Summarization

Abstractive text summarization is a challenging task, and one need to design a mechanism to effectively extract salient information from the source text and then generate a summary. A parsing process of the source text contains critical syntactic or semantic structures, which is useful to generate more accurate summary. However, modeling a parsing tree for text summarization is not trivial due to its non-linear structure and it is harder to deal with a document that includes multiple sentences and their parsing trees. In this paper, we propose to use a graph to connect the parsing trees from the sentences in a document and utilize the stacked graph convolutional networks (GCNs) to learn the syntactic representation for a document. The selective attention mechanism is used to extract salient information in semantic and structural aspect and generate an abstractive summary. We evaluate our approach on the CNN/Daily Mail text summarization dataset. The experimental results show that the proposed GCNs based selective attention approach outperforms the baselines and achieves the state-of-the-art performance on the dataset.

cs.CL↗

DELTA: A DEep learning based Language Technology plAtform

In this paper we present DELTA, a deep learning based language technology platform. DELTA is an end-to-end platform designed to solve industry level natural language and speech processing problems. It integrates most popular neural network models for training as well as comprehensive deployment tools for production. DELTA aims to provide easy and fast experiences for using, deploying, and developing natural language processing and speech models for both academia and industry use cases. We demonstrate the reliable performance with DELTA on several natural language processing and speech tasks, including text classification, named entity recognition, natural language inference, speech recognition, speaker verification, etc. DELTA has been used for developing several state-of-the-art algorithms for publications and delivering real production to serve millions of users.

cs.CL↗

Aperiodic quantum oscillations in the two-dimensional electron gas at the LaAlO3/SrTiO3 interface

Despite several attempts, the intimate electronic structure of two-dimensional electron systems buried at the interface between LaAlO3 and SrTiO3 still remains to be experimentally revealed. Here, we investigate the transport properties of a high-mobility quasi-two-dimensional electron gas at this interface under high magnetic field (55 T) and provide new insights for electronic band structure by analyzing the Shubnikov-de Haas oscillations. Interestingly, the quantum oscillations are not 1/B-periodic and produce a highly non-linear Landau plot (Landau level index versus 1/B). Among possible scenarios, the Roth-Gao-Niu equation provides a natural explanation for 1/B-aperiodic oscillations in relation with the magnetic response functions of the system. Overall, the magneto-transport data are discussed in light of high-resolution scanning transmission electron microscopy analysis of the interface as well as calculations from density functional theory.

cond-mat.mes-hall↗

Collaborative Multi-modal deep learning for the personalized product retrieval in Facebook Marketplace

Facebook Marketplace is quickly gaining momentum among consumers as a favored customer-to-customer (C2C) product trading platform. The recommendation system behind it helps to significantly improve the user experience. Building the recommendation system for Facebook Marketplace is challenging for two reasons: 1) Scalability: the number of products in Facebook Marketplace is huge. Tens of thousands of products need to be scored and recommended within a couple hundred milliseconds for millions of users every day; 2) Cold start: the life span of the C2C products is very short and the user activities on the products are sparse. Thus it is difficult to accumulate enough product level signals for recommendation and we are facing a significant cold start issue. In this paper, we propose to address both the scalability and the cold-start issue by building a collaborative multi-modal deep learning based retrieval system where the compact embeddings for the users and the products are trained with the multi-modal content information. This system shows significant improvement over the benchmark in online and off-line experiments: In the online experiment, it increases the number of messages initiated by the buyer to the seller by +26.95%; in the off-line experiment, it improves the prediction accuracy by +9.58%.

cs.IR↗

The Mechanism of Electrolyte Gating on High-Tc Cuprates: The Role of Oxygen Migration and Electrostatics

Electrolyte gating is widely used to induce large carrier density modulation on solid surfaces to explore various properties. Most of past works have attributed the charge modulation to electrostatic field effect. However, some recent reports have argued that the electrolyte gating effect in VO2, TiO2 and SrTiO3 originated from field-induced oxygen vacancy formation. This gives rise to a controversy about the gating mechanism, and it is therefore vital to reveal the relationship between the role of electrolyte gating and the intrinsic properties of materials. Here, we report entirely different mechanisms of electrolyte gating on two high-Tc cuprates, NdBa2Cu3O7-δ (NBCO) and Pr2-xCexCuO4 (PCCO), with different crystal structures. We show that field-induced oxygen vacancy formation in CuO chains of NBCO plays the dominant role while it is mainly an electrostatic field effect in the case of PCCO. The possible reason is that NBCO has mobile oxygen in CuO chains while PCCO does not. Our study helps clarify the controversy relating to the mechanism of electrolyte gating, leading to a better understanding of the role of oxygen electro migration which is very material specific.

cond-mat.str-el↗

Liquid-Gated High Mobility and Quantum Oscillation of the Two-Dimensional Electron Gas at an Oxide Interface

Electric field effect in electronic double layer transistor (EDLT) configuration with ionic liquids as the dielectric materials is a powerful means of exploring various properties in different materials. Here we demonstrate the modulation of electrical transport properties and extremely high mobility of two-dimensional electron gas at LaAlO$_3$/SrTiO$_3$ (LAO/STO) interface through ionic liquid-assisted electric field effect. By changing the gate voltages, the depletion of charge carrier and the resultant enhancement of electron mobility up to 19380 cm$^2$/Vs are realized, leading to quantum oscillations of the conductivity at the LAO/STO interface. The present results suggest that high-mobility oxide interfaces which exhibit quantum phenomena could be obtained by ionic liquid-assisted field effect.

cond-mat.str-el↗

Measuring Nanoscale Stress Intensity Factors with an Atomic Force Microscope

Atomic Force Microscope images of a crack intersecting the free surface of a glass specimen are taken at different stages of subcritical propagation. From the analysis of image pairs, it is shown that a novel Integrated Digital Image Correlation technique allows to measure stress intensity factors in a quantitative fashion. Image sizes as small as 200 nm can be exploited and the surface displacement fields do not show significant deviations from linear elastic solutions down to a 10 nm distance from the crack tip. Moreover, this analysis gives access to the out-of-plane displacement of the free surface at the crack tip.

physics.class-ph↗