Searcharxiv⌕ Search

arXiv subjects

He He

Publications and source records attributed to He He.

80 records · Page 5Linked to original sources

Delete, Retrieve, Generate: A Simple Approach to Sentiment and Style Transfer

We consider the task of text attribute transfer: transforming a sentence to alter a specific attribute (e.g., sentiment) while preserving its attribute-independent content (e.g., changing "screen is just the right size" to "screen is too small"). Our training data includes only sentences labeled with their attribute (e.g., positive or negative), but not pairs of sentences that differ only in their attributes, so we must learn to disentangle attributes from attribute-independent content in an unsupervised way. Previous work using adversarial methods has struggled to produce high-quality outputs. In this paper, we propose simpler methods motivated by the observation that text attributes are often marked by distinctive phrases (e.g., "too small"). Our strongest method extracts content words by deleting phrases associated with the sentence's original attribute value, retrieves new phrases associated with the target attribute, and uses a neural model to fluently combine these into a final output. On human evaluation, our best method generates grammatical and appropriate responses on 22% more inputs than the best previous system, averaged over three attribute transfer datasets: altering sentiment of reviews on Yelp, altering sentiment of reviews on Amazon, and altering image captions to be more romantic or humorous.

cs.CL↗

Using Sloppy Models for Constrained Emittance Minimization at the Cornell Electron Storage Ring (CESR)

In order to minimize the emittance at the Cornell Electron Storage Ring (CESR), we measure and correct the orbit, dispersion, and transverse coupling of the beam. However, this method is limited by finite measurement resolution of the dispersion, and so a new procedure must be used to further reduce the emittance due to dispersion. In order to achieve this, we use a method based upon the theory of sloppy models. We use a model of the accelerator to create the Hessian matrix which encodes the effects of various corrector magnets on the vertical emittance. A singular value decomposition of this matrix yields the magnet combinations which have the greatest effect on the emittance. We can then adjust these magnet "knobs" sequentially in order to decrease the dispersion and the emittance. We present here comparisons of the effectiveness of this procedure in both experiment and simulation using a variety of CESR lattices. We also discuss techniques to minimize changes to parameters we have already corrected.

physics.acc-ph↗

Learning Symmetric Collaborative Dialogue Agents with Dynamic Knowledge Graph Embeddings

We study a symmetric collaborative dialogue setting in which two agents, each with private knowledge, must strategically communicate to achieve a common goal. The open-ended dialogue state in this setting poses new challenges for existing dialogue systems. We collected a dataset of 11K human-human dialogues, which exhibits interesting lexical, semantic, and strategic elements. To model both structured knowledge and unstructured language, we propose a neural model with dynamic knowledge graph embeddings that evolve as the dialogue progresses. Automatic and human evaluations show that our model is both more effective at achieving the goal and more human-like than baseline neural and rule-based models.

cs.CL↗

Opponent Modeling in Deep Reinforcement Learning

Opponent modeling is necessary in multi-agent settings where secondary agents with competing goals also adapt their strategies, yet it remains challenging because strategies interact with each other and change. Most previous work focuses on developing probabilistic models or parameterized strategies for specific applications. Inspired by the recent success of deep reinforcement learning, we present neural-based models that jointly learn a policy and the behavior of opponents. Instead of explicitly predicting the opponent's action, we encode observation of the opponents into a deep Q-Network (DQN); however, we retain explicit modeling (if desired) using multitasking. By using a Mixture-of-Experts architecture, our model automatically discovers different strategy patterns of opponents without extra supervision. We evaluate our models on a simulated soccer game and a popular trivia game, showing superior performance over DQN and its variants.

cs.LG↗

A Credit Assignment Compiler for Joint Prediction

Many machine learning applications involve jointly predicting multiple mutually dependent output variables. Learning to search is a family of methods where the complex decision problem is cast into a sequence of decisions via a search space. Although these methods have shown promise both in theory and in practice, implementing them has been burdensomely awkward. In this paper, we show the search space can be defined by an arbitrary imperative program, turning learning to search into a credit assignment compiler. Altogether with the algorithmic improvements for the compiler, we radically reduce the complexity of programming and the running time. We demonstrate the feasibility of our approach on multiple joint prediction tasks. In all cases, we obtain accuracies as high as alternative approaches, at drastically reduced execution and programming time.

cs.LG↗

Active Information Acquisition

We propose a general framework for sequential and dynamic acquisition of useful information in order to solve a particular task. While our goal could in principle be tackled by general reinforcement learning, our particular setting is constrained enough to allow more efficient algorithms. In this paper, we work under the Learning to Search framework and show how to formulate the goal of finding a dynamic information acquisition policy in that framework. We apply our formulation on two tasks, sentiment analysis and image recognition, and show that the learned policies exhibit good statistical performance. As an emergent byproduct, the learned policies show a tendency to focus on the most prominent parts of each instance and give harder instances more attention without explicitly being trained to do so.

stat.ML↗

Learning to Search for Dependencies

We demonstrate that a dependency parser can be built using a credit assignment compiler which removes the burden of worrying about low-level machine learning details from the parser implementation. The result is a simple parser which robustly applies to many languages that provides similar statistical and computational performance with best-to-date transition-based parsing approaches, while avoiding various downsides including randomization, extra feature requirements, and custom learning algorithms.

cs.CL↗

Epitaxial strain induced magnetic transitions and phonon instabilities of the tetragonal SrRuO3

Using density-functional theory calculations, we investigate the magnetic as well as the dynamical properties of tetragonal SrRuO3 (SRO) under the influence of epitaxial strain. It is found that both the tensile and compressive strain in the xy-plane could induce the abrupt change in the magnetic moment of Ru atom. In particular, under the in-plane ~4% compressive strain, a ferromagnetic to nonmagnetic transition is induced. Whereas for the tensile strain larger than 3%, the Ru magnetic moment drops gradually with the increase of the strain, exhibiting a weak ferromagnetic state. We find that such magnetic transitions could be qualitatively explained by the Stoner model. In addition, frozen phonon calculations at Γ point reveal structural instabilities could occur under both compressive and tensile strains. Such instabilities are very similar to those of the ferroelectric perovskite oxides, even though SRO remains to be metallic in the range we studied. These might have influence on the physical properties of oxide supercells taking SRO as constituent.

cond-mat.str-el↗