Searcharxiv⌕ Search

arXiv subjects

John Hewitt

Publications and source records attributed to John Hewitt.

27 records · Page 2Linked to original sources

Probing artificial neural networks: insights from neuroscience

A major challenge in both neuroscience and machine learning is the development of useful tools for understanding complex information processing systems. One such tool is probes, i.e., supervised models that relate features of interest to activation patterns arising in biological or artificial neural networks. Neuroscience has paved the way in using such models through numerous studies conducted in recent decades. In this work, we draw insights from neuroscience to help guide probing research in machine learning. We highlight two important design choices for probes $-$ direction and expressivity $-$ and relate these choices to research goals. We argue that specific research goals play a paramount role when designing a probe and encourage future probing studies to be explicit in stating these goals.

cs.LG↗

RNNs can generate bounded hierarchical languages with optimal memory

Recurrent neural networks empirically generate natural language with high syntactic fidelity. However, their success is not well-understood theoretically. We provide theoretical insight into this success, proving in a finite-precision setting that RNNs can efficiently generate bounded hierarchical languages that reflect the scaffolding of natural language syntax. We introduce Dyck-($k$,$m$), the language of well-nested brackets (of $k$ types) and $m$-bounded nesting depth, reflecting the bounded memory needs and long-distance dependencies of natural language syntax. The best known results use $O(k^{\frac{m}{2}})$ memory (hidden units) to generate these languages. We prove that an RNN with $O(m \log k)$ hidden units suffices, an exponential reduction in memory, by an explicit construction. Finally, we show that no algorithm, even with unbounded computation, can suffice with $o(m \log k)$ hidden units.

cs.CL↗

The EOS Decision and Length Extrapolation

Extrapolation to unseen sequence lengths is a challenge for neural generative models of language. In this work, we characterize the effect on length extrapolation of a modeling decision often overlooked: predicting the end of the generative process through the use of a special end-of-sequence (EOS) vocabulary item. We study an oracle setting - forcing models to generate to the correct sequence length at test time - to compare the length-extrapolative behavior of networks trained to predict EOS (+EOS) with networks not trained to (-EOS). We find that -EOS substantially outperforms +EOS, for example extrapolating well to lengths 10 times longer than those seen at training time in a bracket closing task, as well as achieving a 40% improvement over +EOS in the difficult SCAN dataset length generalization task. By comparing the hidden states and dynamics of -EOS and +EOS models, we observe that +EOS models fail to generalize because they (1) unnecessarily stratify their hidden states by their linear position is a sequence (structures we call length manifolds) or (2) get stuck in clusters (which we refer to as length attractors) once the EOS token is the highest-probability prediction.

cs.CL↗

Finding Universal Grammatical Relations in Multilingual BERT

Recent work has found evidence that Multilingual BERT (mBERT), a transformer-based multilingual masked language model, is capable of zero-shot cross-lingual transfer, suggesting that some aspects of its representations are shared cross-lingually. To better understand this overlap, we extend recent work on finding syntactic trees in neural networks' internal representations to the multilingual setting. We show that subspaces of mBERT representations recover syntactic tree distances in languages other than English, and that these subspaces are approximately shared across languages. Motivated by these results, we present an unsupervised analysis method that provides evidence mBERT learns representations of syntactic dependency labels, in the form of clusters which largely agree with the Universal Dependencies taxonomy. This evidence suggests that even without explicit supervision, multilingual masked language models learn certain linguistic universals.

cs.CL↗

Designing and Interpreting Probes with Control Tasks

Probes, supervised models trained to predict properties (like parts-of-speech) from representations (like ELMo), have achieved high accuracy on a range of linguistic tasks. But does this mean that the representations encode linguistic structure or just that the probe has learned the linguistic task? In this paper, we propose control tasks, which associate word types with random outputs, to complement linguistic tasks. By construction, these tasks can only be learned by the probe itself. So a good probe, (one that reflects the representation), should be selective, achieving high linguistic task accuracy and low control task accuracy. The selectivity of a probe puts linguistic task accuracy in context with the probe's capacity to memorize from word types. We construct control tasks for English part-of-speech tagging and dependency edge prediction, and show that popular probes on ELMo representations are not selective. We also find that dropout, commonly used to control probe complexity, is ineffective for improving selectivity of MLPs, but that other forms of regularization are effective. Finally, we find that while probes on the first layer of ELMo yield slightly better part-of-speech tagging accuracy than the second, probes on the second layer are substantially more selective, which raises the question of which layer better represents parts-of-speech.

cs.CL↗

Simple, Fast, Accurate Intent Classification and Slot Labeling for Goal-Oriented Dialogue Systems

With the advent of conversational assistants, like Amazon Alexa, Google Now, etc., dialogue systems are gaining a lot of traction, especially in industrial setting. These systems typically consist of Spoken Language understanding component which, in turn, consists of two tasks - Intent Classification (IC) and Slot Labeling (SL). Generally, these two tasks are modeled together jointly to achieve best performance. However, this joint modeling adds to model obfuscation. In this work, we first design framework for a modularization of joint IC-SL task to enhance architecture transparency. Then, we explore a number of self-attention, convolutional, and recurrent models, contributing a large-scale analysis of modeling paradigms for IC+SL across two datasets. Finally, using this framework, we propose a class of 'label-recurrent' models that otherwise non-recurrent, with a 10-dimensional representation of the label history, and show that our proposed systems are easy to interpret, highly accurate (achieving over 30% error reduction in SL over the state-of-the-art on the Snips dataset), as well as fast, at 2x the inference and 2/3 to 1/2 the training time of comparable recurrent models, thus giving an edge in critical real-world systems.

cs.CL↗

XNMT: The eXtensible Neural Machine Translation Toolkit

This paper describes XNMT, the eXtensible Neural Machine Translation toolkit. XNMT distin- guishes itself from other open-source NMT toolkits by its focus on modular code design, with the purpose of enabling fast iteration in research and replicable, reliable results. In this paper we describe the design of XNMT and its experiment configuration system, and demonstrate its utility on the tasks of machine translation, speech recognition, and multi-tasked machine translation/parsing. XNMT is available open-source at https://github.com/neulab/xnmt

cs.CL↗

Discovery of X-ray Emission from the Galactic Supernova Remnant G32.8-0.1 with Suzaku

We present the first dedicated X-ray study of the supernova remnant (SNR) G32.8-0.1 (Kes 78) with Suzaku. X-ray emission from the whole SNR shell has been detected for the first time. The X-ray morphology is well correlated with the emission from the radio shell, while anti-correlated with the molecular cloud found in the SNR field. The X-ray spectrum shows not only conventional low-temperature (kT ~ 0.6 keV) thermal emission in a non-equilibrium ionization state, but also a very high temperature (kT ~ 3.4 keV) component with a very low ionization timescale (~ 2.7e9 cm^{-3}s), or a hard non-thermal component with a photon index Gamma~2.3. The average density of the low-temperature plasma is rather low, of the order of 10^{-3}--10^{-2} cm^{-3}, implying that this SNR is expanding into a low-density cavity. We discuss the X-ray emission of the SNR, also detected in TeV with H.E.S.S., together with multi-wavelength studies of the remnant and other gamma-ray emitting SNRs, such as W28 and RCW 86. Analysis of a time-variable source, 2XMM J185114.3-000004, found in the northern part of the SNR, is also reported for the first time. Rapid time variability and a heavily absorbed hard X-ray spectrum suggest that this source could be a new supergiant fast X-ray transient.

astro-ph.HE↗

New Identification of the Mixed-Morphology Supernova Remnant G298.6-0.0 with Possible Gamma-ray Association

We present an X-ray analysis on the Galactic supernova remnant (SNR) G298.6-0.0 with Suzaku. The X-ray image shows a center-filled structure inside the radio shell, implying this SNR is categorized as a mixed-morphology (MM) SNR. The spectrum is well reproduced by a single temperature plasma model in ionization equilibrium, with a temperature of 0.78 (0.70-0.87) keV. The total plasma mass of 30 solar mass indicates that the plasma has interstellar medium origin. The association with a GeV gamma-ray source 3FGL J1214.0-6236 on the shell of the SNR is discussed, in comparison with other MM SNRs with GeV gamma-ray associations. It is found that the flux ratio between absorption-corrected thermal X-rays and GeV gamma-rays decreases as the MM SNRs evolve to larger physical sizes. The absorption-corrected X-ray flux of G298.6-0.0 and the GeV gamma-ray flux of 3FGL J1214.0-6236 closely follow this trend, implying that 3FGL J1214.0-6236 is likely to be the GeV counterpart of G298.6-0.0.

astro-ph.HE↗