Searcharxiv⌕ Search

arXiv subjects

Fei Cheng

Publications and source records attributed to Fei Cheng.

71 records · Page 4Linked to original sources

Seeking Diverse Reasoning Logic: Controlled Equation Expression Generation for Solving Math Word Problems

To solve Math Word Problems, human students leverage diverse reasoning logic that reaches different possible equation solutions. However, the mainstream sequence-to-sequence approach of automatic solvers aims to decode a fixed solution equation supervised by human annotation. In this paper, we propose a controlled equation generation solver by leveraging a set of control codes to guide the model to consider certain reasoning logic and decode the corresponding equations expressions transformed from the human reference. The empirical results suggest that our method universally improves the performance on single-unknown (Math23K) and multiple-unknown (DRAW1K, HMWP) benchmarks, with substantial improvements up to 13.2% accuracy on the challenging multiple-unknown datasets.

cs.CL↗

Textual Enhanced Contrastive Learning for Solving Math Word Problems

Solving math word problems is the task that analyses the relation of quantities and requires an accurate understanding of contextual natural language information. Recent studies show that current models rely on shallow heuristics to predict solutions and could be easily misled by small textual perturbations. To address this problem, we propose a Textual Enhanced Contrastive Learning framework, which enforces the models to distinguish semantically similar examples while holding different mathematical logic. We adopt a self-supervised manner strategy to enrich examples with subtle textual variance by textual reordering or problem re-construction. We then retrieve the hardest to differentiate samples from both equation and textual perspectives and guide the model to learn their representations. Experimental results show that our method achieves state-of-the-art on both widely used benchmark datasets and also exquisitely designed challenge datasets in English and Chinese. \footnote{Our code and data is available at \url{https://github.com/yiyunya/Textual_CL_MWP}

cs.CL↗

Reverse Operation based Data Augmentation for Solving Math Word Problems

Automatically solving math word problems is a critical task in the field of natural language processing. Recent models have reached their performance bottleneck and require more high-quality data for training. We propose a novel data augmentation method that reverses the mathematical logic of math word problems to produce new high-quality math problems and introduce new knowledge points that can benefit learning the mathematical reasoning logic. We apply the augmented data on two SOTA math word problem solving models and compare our results with a strong data augmentation baseline. Experimental results show the effectiveness of our approach. We release our code and data at https://github.com/yiyunya/RODA.

cs.CL↗

Cross-lingual Adaption Model-Agnostic Meta-Learning for Natural Language Understanding

Meta learning with auxiliary languages has demonstrated promising improvements for cross-lingual natural language processing. However, previous studies sample the meta-training and meta-testing data from the same language, which limits the ability of the model for cross-lingual transfer. In this paper, we propose XLA-MAML, which performs direct cross-lingual adaption in the meta-learning stage. We conduct zero-shot and few-shot experiments on Natural Language Inference and Question Answering. The experimental results demonstrate the effectiveness of our method across different languages, tasks, and pretrained models. We also give analysis on various cross-lingual specific settings for meta-learning including sampling strategy and parallelism.

cs.CL↗

JaMIE: A Pipeline Japanese Medical Information Extraction System

We present an open-access natural language processing toolkit for Japanese medical information extraction. We first propose a novel relation annotation schema for investigating the medical and temporal relations between medical entities in Japanese medical reports. We experiment with the practical annotation scenarios by separately annotating two different types of reports. We design a pipeline system with three components for recognizing medical entities, classifying entity modalities, and extracting relations. The empirical results show accurate analyzing performance and suggest the satisfactory annotation quality, the effective annotation strategy for targeting report types, and the superiority of the latest contextual embedding models.

cs.CL↗

OCHADAI-KYOTO at SemEval-2021 Task 1: Enhancing Model Generalization and Robustness for Lexical Complexity Prediction

We propose an ensemble model for predicting the lexical complexity of words and multiword expressions (MWEs). The model receives as input a sentence with a target word or MWEand outputs its complexity score. Given that a key challenge with this task is the limited size of annotated data, our model relies on pretrained contextual representations from different state-of-the-art transformer-based language models (i.e., BERT and RoBERTa), and on a variety of training methods for further enhancing model generalization and robustness:multi-step fine-tuning and multi-task learning, and adversarial training. Additionally, we propose to enrich contextual representations by adding hand-crafted features during training. Our model achieved competitive results and ranked among the top-10 systems in both sub-tasks.

cs.CL↗

ShipSRDet: An End-to-End Remote Sensing Ship Detector Using Super-Resolved Feature Representation

High-resolution remote sensing images can provide abundant appearance information for ship detection. Although several existing methods use image super-resolution (SR) approaches to improve the detection performance, they consider image SR and ship detection as two separate processes and overlook the internal coherence between these two correlated tasks. In this paper, we explore the potential benefits introduced by image SR to ship detection, and propose an end-to-end network named ShipSRDet. In our method, we not only feed the super-resolved images to the detector but also integrate the intermediate features of the SR network with those of the detection network. In this way, the informative feature representation extracted by the SR network can be fully used for ship detection. Experimental results on the HRSC dataset validate the effectiveness of our method. Our ShipSRDet can recover the missing details from the input image and achieves promising ship detection performance.

cs.CV↗

A Hybrid Bandit Framework for Diversified Recommendation

The interactive recommender systems involve users in the recommendation procedure by receiving timely user feedback to update the recommendation policy. Therefore, they are widely used in real application scenarios. Previous interactive recommendation methods primarily focus on learning users' personalized preferences on the relevance properties of an item set. However, the investigation of users' personalized preferences on the diversity properties of an item set is usually ignored. To overcome this problem, we propose the Linear Modular Dispersion Bandit (LMDB) framework, which is an online learning setting for optimizing a combination of modular functions and dispersion functions. Specifically, LMDB employs modular functions to model the relevance properties of each item, and dispersion functions to describe the diversity properties of an item set. Moreover, we also develop a learning algorithm, called Linear Modular Dispersion Hybrid (LMDH) to solve the LMDB problem and derive a gap-free bound on its n-step regret. Extensive experiments on real datasets are performed to demonstrate the effectiveness of the proposed LMDB framework in balancing the recommendation accuracy and diversity.

cs.IR↗

A System for Worldwide COVID-19 Information Aggregation

The global pandemic of COVID-19 has made the public pay close attention to related news, covering various domains, such as sanitation, treatment, and effects on education. Meanwhile, the COVID-19 condition is very different among the countries (e.g., policies and development of the epidemic), and thus citizens would be interested in news in foreign countries. We build a system for worldwide COVID-19 information aggregation containing reliable articles from 10 regions in 7 languages sorted by topics. Our reliable COVID-19 related website dataset collected through crowdsourcing ensures the quality of the articles. A neural machine translation module translates articles in other languages into Japanese and English. A BERT-based topic-classifier trained on our article-topic pair dataset helps users find their interested information efficiently by putting articles into different categories.

cs.CL↗

Minimize Exposure Bias of Seq2Seq Models in Joint Entity and Relation Extraction

Joint entity and relation extraction aims to extract relation triplets from plain text directly. Prior work leverages Sequence-to-Sequence (Seq2Seq) models for triplet sequence generation. However, Seq2Seq enforces an unnecessary order on the unordered triplets and involves a large decoding length associated with error accumulation. These introduce exposure bias, which may cause the models overfit to the frequent label combination, thus deteriorating the generalization. We propose a novel Sequence-to-Unordered-Multi-Tree (Seq2UMTree) model to minimize the effects of exposure bias by limiting the decoding length to three within a triplet and removing the order among triplets. We evaluate our model on two datasets, DuIE and NYT, and systematically study how exposure bias alters the performance of Seq2Seq models. Experiments show that the state-of-the-art Seq2Seq model overfits to both datasets while Seq2UMTree shows significantly better generalization. Our code is available at https://github.com/WindChimeRan/OpenJERE .

cs.CL↗

Epitaxial Growth of Two-dimensional Insulator Monolayer Honeycomb BeO

The emergence of two-dimensional (2D) materials launched a fascinating frontier of flatland electronics. Most crystalline atomic layer materials are based on layered van der Waals materials with weak interlayer bonding, which naturally leads to thermodynamically stable monolayers. We report the synthesis of a 2D insulator comprised of a single atomic sheet of honeycomb structure BeO (h-BeO), although its bulk counterpart has a wurtzite structure. The h-BeO is grown by molecular beam epitaxy (MBE) on Ag(111) thin films that are conveniently grown on Si(111) wafers. Using scanning tunneling microscopy and spectroscopy (STM/S), the honeycomb BeO lattice constant is determined to be 2.65 angstrom with an insulating band gap of 6 eV. Our low energy electron diffraction (LEED) measurements indicate that the h-BeO forms a continuous layer with good crystallinity at the millimeter scale. Moiré pattern analysis shows the BeO honeycomb structure maintains long range phase coherence in atomic registry even across Ag steps. We find that the interaction between the h-BeO layer and the Ag(111) substrate is weak by using STS and complimentary density functional theory calculations. We not only demonstrate the feasibility of growing h-BeO monolayers by MBE, but also illustrate that the large-scale growth, weak substrate interactions, and long-range crystallinity make h-BeO an attractive candidate for future technological applications. More significantly, the ability to create a stable single crystalline atomic sheet without a bulk layered counterpart is an intriguing approach to tailoring novel 2D electronic materials.

cond-mat.mes-hall↗

Predicting Event Time by Classifying Sub-Level Temporal Relations Induced from a Unified Representation of Time Anchors

Extracting event time from news articles is a challenging but attractive task. In contrast to the most existing pair-wised temporal link annotation, Reimers et al.(2016) proposed to annotate the time anchor (a.k.a. the exact time) of each event. Their work represents time anchors with discrete representations of Single-Day/Multi-Day and Certain/Uncertain. This increases the complexity of modeling the temporal relations between two time anchors, which cannot be categorized into the relations of Allen's interval algebra (Allen, 1990). In this paper, we propose an effective method to decompose such complex temporal relations into sub-level relations by introducing a unified quadruple representation for both Single-Day/Multi-Day and Certain/Uncertain time anchors. The temporal relation classifiers are trained in a multi-label classification manner. The system structure of our approach is much simpler than the existing decision tree model (Reimers et al., 2018), which is composed by a dozen of node classifiers. Another contribution of this work is to construct a larger event time corpus (256 news documents) with a reasonable Inter-Annotator Agreement (IAA), for the purpose of overcoming the data shortage of the existing event time corpus (36 news documents). The empirical results show our approach outperforms the state-of-the-art decision tree model and the increase of data size obtained a significant improvement of performance.

cs.CL↗

Adversarial Training for Commonsense Inference

We propose an AdversariaL training algorithm for commonsense InferenCE (ALICE). We apply small perturbations to word embeddings and minimize the resultant adversarial risk to regularize the model. We exploit a novel combination of two different approaches to estimate these perturbations: 1) using the true label and 2) using the model prediction. Without relying on any human-crafted features, knowledge bases, or additional datasets other than the target datasets, our model boosts the fine-tuning performance of RoBERTa, achieving competitive results on multiple reading comprehension datasets that require commonsense inference.

cs.CL↗

Pre-training via Leveraging Assisting Languages and Data Selection for Neural Machine Translation

Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks in low-resource settings. However, large monolingual corpora might not always be available for the languages of interest (LOI). To this end, we propose to exploit monolingual corpora of other languages to complement the scarcity of monolingual corpora for the LOI. A case study of low-resource Japanese-English neural machine translation (NMT) reveals that leveraging large Chinese and French monolingual corpora can help overcome the shortage of Japanese and English monolingual corpora, respectively, for S2S pre-training. We further show how to utilize script mapping (Chinese to Japanese) to increase the similarity between the two monolingual corpora leading to further improvements in translation quality. Additionally, we propose simple data-selection techniques to be used prior to pre-training that significantly impact the quality of S2S pre-training. An empirical comparison of our proposed methods reveals that leveraging assisting language monolingual corpora, data selection and script mapping are extremely important for NMT pre-training in low-resource scenarios.

cs.CL↗

Random Occlusion-recovery for Person Re-identification

As a basic task of multi-camera surveillance system, person re-identification aims to re-identify a query pedestrian observed from non-overlapping multiple cameras or across different time with a single camera. Recently, deep learning-based person re-identification models have achieved great success in many benchmarks. However, these supervised models require a large amount of labeled image data, and the process of manual labeling spends much manpower and time. In this study, we introduce a method to automatically synthesize labeled person images and adopt them to increase the sample number per identity for person re-identification datasets. To be specific, we use block rectangles to randomly occlude pedestrian images. Then, a generative adversarial network (GAN) model is proposed to use paired occluded and original images to synthesize the de-occluded images that similar but not identical to the original image. Afterwards, we annotate the de-occluded images with the same labels of their corresponding raw images and use them to augment the number of samples per identity. Finally, we use the augmented datasets to train baseline model. The experiment results on CUHK03, Market-1501 and DukeMTMC-reID datasets show that the effectiveness of the proposed method.

cs.CV↗

Single Crystalline Silver Films for Plasmonics: From Monolayer to Optically Thick Film

Epitaxial growth of single crystalline noble metals on dielectric substrates has received tremendous attention recently due to their technological potentials as low loss plasmonic materials. Currently there are two different growth approaches, each with its strengths and weaknesses. One adopts a sophisticated molecular beam epitaxial procedure to grow atomically smooth epitaxial Ag films. However, the procedure is rather slow and becomes impractical to grow films with thickness > 50 nm. Another approach adopts a growth process using rapid e-beam deposition which is capable of growing single crystalline Ag films in the thick regime (> 300 nm). However, the rapid growth procedure makes it difficult to control film thickness precisely, i.e., the method is not applicable to growing thin epitaxial films. Here we report a universal approach to grow atomically smooth epitaxial Ag films with precise thickness control from a few monolayers to the optically thick regime, overcoming the limitations of the two aforementioned methods. In addition, we develop an in-situ growth of aluminum oxide as the capping layer which exhibits excellent properties protecting the epitaxial Ag films. The performance of the epitaxial Ag films as a function of the film thickness is investigated by directly measuring the propagation length of the surface plasmon polaritons (SPPs) as well as their device performance to support a waveguide plasmonic nanolaser in infrared incorporating an InGaAsP quantum well as the gain media.

cond-mat.mtrl-sci↗

Tailoring Semiconductor Lateral Multi-junctions for Giant Photoconductivity Enhancement

Semiconductor heterostructures have played a critical role as the enabler for new science and technology. The emergence of transition metal dichalcogenides (TMDs) as atomically thin semiconductors has opened new frontiers in semiconductor heterostructures either by stacking different TMDs to form vertical heterojunctions or by stitching them laterally to form lateral heterojunctions via direct growth. In conventional semiconductor heterostructures, the design of multi-junctions is critical to achieve carrier confinement. Analogously, we report successful synthesis of monolayer WS2/WS2(1-x)Se2x/WS2 multi-junction lateral heterostructure via direct growth by chemical vapor deposition. The grown structures are characterized by Raman, photoluminescence, and annular dark-field scanning transmission electron microscopy to determine its lateral compositional profile. More importantly, using microwave impedance microscopy, we demonstrate that the local photoconductivity in the alloy region can be tailored and enhanced by 2 orders of magnitude over pure WS2. Finite element analysis confirms that this effect is due to the carrier diffusion and confinement into the alloy region. Our work exemplifies the technological potential of atomically thin lateral heterostructures in optoelectronic applications.

cond-mat.mtrl-sci↗