Searcharxiv⌕ Search

arXiv subjects

Christopher Clark

Publications and source records attributed to Christopher Clark.

33 records · Page 2Linked to original sources

Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text

Communicating with humans is challenging for AIs because it requires a shared understanding of the world, complex semantics (e.g., metaphors or analogies), and at times multi-modal gestures (e.g., pointing with a finger, or an arrow in a diagram). We investigate these challenges in the context of Iconary, a collaborative game of drawing and guessing based on Pictionary, that poses a novel challenge for the research community. In Iconary, a Guesser tries to identify a phrase that a Drawer is drawing by composing icons, and the Drawer iteratively revises the drawing to help the Guesser in response. This back-and-forth often uses canonical scenes, visual metaphor, or icon compositions to express challenging words, making it an ideal test for mixing language and visual/symbolic communication in AI. We propose models to play Iconary and train them on over 55,000 games between human players. Our models are skillful players and are able to employ world knowledge in language models to play with words unseen during training. Elite human players outperform our models, particularly at the drawing task, leaving an important gap for future research to address. We release our dataset, code, and evaluation setup as a challenge to the community at http://www.github.com/allenai/iconary.

cs.CL↗

Learning to Model and Ignore Dataset Bias with Mixed Capacity Ensembles

Many datasets have been shown to contain incidental correlations created by idiosyncrasies in the data collection process. For example, sentence entailment datasets can have spurious word-class correlations if nearly all contradiction sentences contain the word "not", and image recognition datasets can have tell-tale object-background correlations if dogs are always indoors. In this paper, we propose a method that can automatically detect and ignore these kinds of dataset-specific patterns, which we call dataset biases. Our method trains a lower capacity model in an ensemble with a higher capacity model. During training, the lower capacity model learns to capture relatively shallow correlations, which we hypothesize are likely to reflect dataset bias. This frees the higher capacity model to focus on patterns that should generalize better. We ensure the models learn non-overlapping approaches by introducing a novel method to make them conditionally independent. Importantly, our approach does not require the bias to be known in advance. We evaluate performance on synthetic datasets, and four datasets built to penalize models that exploit known biases on textual entailment, visual question answering, and image recognition tasks. We show improvement in all settings, including a 10 point gain on the visual question answering dataset.

cs.LG↗

Long-distance Detection of Bioacoustic Events with Per-channel Energy Normalization

This paper proposes to perform unsupervised detection of bioacoustic events by pooling the magnitudes of spectrogram frames after per-channel energy normalization (PCEN). Although PCEN was originally developed for speech recognition, it also has beneficial effects in enhancing animal vocalizations, despite the presence of atmospheric absorption and intermittent noise. We prove that PCEN generalizes logarithm-based spectral flux, yet with a tunable time scale for background noise estimation. In comparison with pointwise logarithm, PCEN reduces false alarm rate by 50x in the near field and 5x in the far field, both on avian and marine bioacoustic datasets. Such improvements come at moderate computational cost and require no human intervention, thus heralding a promising future for PCEN in bioacoustics.

cs.SD↗

Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases

State-of-the-art models often make use of superficial patterns in the data that do not generalize well to out-of-domain or adversarial settings. For example, textual entailment models often learn that particular key words imply entailment, irrespective of context, and visual question answering models learn to predict prototypical answers, without considering evidence in the image. In this paper, we show that if we have prior knowledge of such biases, we can train a model to be more robust to domain shift. Our method has two stages: we (1) train a naive model that makes predictions exclusively based on dataset biases, and (2) train a robust model as part of an ensemble with the naive one in order to encourage it to focus on other patterns in the data that are more likely to generalize. Experiments on five datasets with out-of-domain test sets show significantly improved robustness in all settings, including a 12 point gain on a changing priors visual question answering dataset and a 9 point gain on an adversarial question answering test set.

cs.CL↗

BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

In this paper we study yes/no questions that are naturally occurring --- meaning that they are generated in unprompted and unconstrained settings. We build a reading comprehension dataset, BoolQ, of such questions, and show that they are unexpectedly challenging. They often query for complex, non-factoid information, and require difficult entailment-like inference to solve. We also explore the effectiveness of a range of transfer learning baselines. We find that transferring from entailment data is more effective than transferring from paraphrase or extractive QA data, and that it, surprisingly, continues to be very beneficial even when starting from massive pre-trained language models such as BERT. Our best method trains BERT on MultiNLI and then re-trains it on our train set. It achieves 80.4% accuracy compared to 90% accuracy of human annotators (and 62% majority-baseline), leaving a significant gap for future work.

cs.CL↗

The Life Cycle of Dust

Dust offers a unique probe of the interstellar medium (ISM) across multiple size, density, and temperature scales. Dust is detected in outflows of evolved stars, star-forming molecular clouds, planet-forming disks, and even in galaxies at the dawn of the Universe. These grains also have a profound effect on various astrophysical phenomena from thermal balance and extinction in galaxies to the building blocks for planets, and changes in dust grain properties will affect all of these phenomena. A full understanding of dust in all of its forms and stages requires a multi-disciplinary investigation of the dust life cycle. Such an investigation can be achieved with a statistical study of dust properties across stellar evolution, star and planet formation, and redshift. Current and future instrumentation will enable this investigation through fast and sensitive observations in dust continuum, polarization, and spectroscopy from near-infrared to millimeter wavelengths.

astro-ph.GA↗

Interstellar Dust Grains: Ultraviolet and Mid-IR Extinction Curves

Interstellar dust plays a central role in shaping the detailed structure of the interstellar medium, thus strongly influencing star formation and galaxy evolution. Dust extinction provides one of the main pillars of our understanding of interstellar dust while also often being one of the limiting factors when interpreting observations of distant objects, including resolved and unresolved galaxies. The ultraviolet (UV) and mid-infrared (MIR) wavelength regimes exhibit features of the main components of dust, carbonaceous and silicate materials, and therefore provide the most fruitful avenue for detailed extinction curve studies. Our current picture of extinction curves is strongly biased to nearby regions in the Milky Way. The small number of UV extinction curves measured in the Local Group (mainly Magellanic Clouds) clearly indicates that the range of dust properties is significantly broader than those inferred from the UV extinction characteristics of local regions of the Milky Way. Obtaining statistically significant samples of UV and MIR extinction measurements for all the dusty Local Group galaxies will provide, for the first time, a basis for understanding dust grains over a wide range of environments. Obtaining such observations requires sensitive medium-band UV, blue-optical, and mid-IR imaging and followup R ~ 1000 spectroscopy of thousands of sources. Such a census will revolutionize our understanding of the dependence of dust properties on local environment providing both an empirical description of the effects of dust on observations as well as strong constraints on dust grain and evolution models.

astro-ph.GA↗

Astro2020: Unleashing the Potential of Dust Emission as a Window onto Galaxy Evolution

We present the severe, systematic uncertainties currently facing our understanding of dust emission, which stymie our ability to truly exploit dust as a tool for studying galaxy evolution. We propose a program of study to tackle these uncertainties, describe the necessary facilities, and discuss the potential science gains that will result. This white paper was submitted to the US National Academies' Astro2020 Decadal Survey on Astronomy and Astrophysics.

astro-ph.GA↗

The Causes of the Red Sequence, the Blue Cloud, the Green Valley and the Green Mountain

The galaxies found in optical surveys fall in two distinct regions of a diagram of optical colour versus absolute magnitude: the red sequence and the blue cloud with the green valley in between. We show that the galaxies found in a submillimetre survey have almost the opposite distribution in this diagram, forming a `green mountain'. We show that these distinctive distributions follow naturally from a single, continuous, curved Galaxy Sequence in a diagram of specific star-formation rate versus stellar mass without there being the need for a separate star-forming galaxy Main Sequence and region of passive galaxies. The cause of the red sequence and the blue cloud is the geometric mapping between stellar mass/specific star-formation rate and absolute magnitude/colour, which distorts a continuous Galaxy Sequence in the diagram of intrinsic properties into a bimodal distribution in the diagram of observed properties. The cause of the green mountain is Malmquist bias in the submillimetre waveband, with submillimetre surveys tending to select galaxies on the curve of the Galaxy Sequence, which have the highest ratios of submillimetre-to-optical luminosity. This effect, working in reverse, causes galaxies on the curve of the Galaxy Sequence to be underrepresented in optical samples, deepening the green valley. The green valley is therefore not evidence (1) for there being two distinct populations of galaxies, (2) for galaxies in this region evolving more quickly than galaxies in the blue cloud and the red sequence, (c) for rapid quenching processes in the galaxy population.

astro-ph.GA↗

Deep contextualized word representations

We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.g., syntax and semantics), and (2) how these uses vary across linguistic contexts (i.e., to model polysemy). Our word vectors are learned functions of the internal states of a deep bidirectional language model (biLM), which is pre-trained on a large text corpus. We show that these representations can be easily added to existing models and significantly improve the state of the art across six challenging NLP problems, including question answering, textual entailment and sentiment analysis. We also present an analysis showing that exposing the deep internals of the pre-trained network is crucial, allowing downstream models to mix different types of semi-supervision signals.

cs.CL↗

Simple and Effective Multi-Paragraph Reading Comprehension

We consider the problem of adapting neural paragraph-level question answering models to the case where entire documents are given as input. Our proposed solution trains models to produce well calibrated confidence scores for their results on individual paragraphs. We sample multiple paragraphs from the documents during training, and use a shared-normalization training objective that encourages the model to produce globally correct output. We combine this method with a state-of-the-art pipeline for training models on document QA data. Experiments demonstrate strong performance on several document QA datasets. Overall, we are able to achieve a score of 71.3 F1 on the web portion of TriviaQA, a large improvement from the 56.7 F1 of the previous best system.

cs.CL↗

The Astropy Problem

The Astropy Project (http://astropy.org) is, in its own words, "a community effort to develop a single core package for Astronomy in Python and foster interoperability between Python astronomy packages." For five years this project has been managed, written, and operated as a grassroots, self-organized, almost entirely volunteer effort while the software is used by the majority of the astronomical community. Despite this, the project has always been and remains to this day effectively unfunded. Further, contributors receive little or no formal recognition for creating and supporting what is now critical software. This paper explores the problem in detail, outlines possible solutions to correct this, and presents a few suggestions on how to address the sustainability of general purpose astronomical software.

astro-ph.IM↗

Teaching Deep Convolutional Neural Networks to Play Go

Mastering the game of Go has remained a long standing challenge to the field of AI. Modern computer Go systems rely on processing millions of possible future positions to play well, but intuitively a stronger and more 'humanlike' way to play the game would be to rely on pattern recognition abilities rather then brute force computation. Following this sentiment, we train deep convolutional neural networks to play Go by training them to predict the moves made by expert Go players. To solve this problem we introduce a number of novel techniques, including a method of tying weights in the network to 'hard code' symmetries that are expect to exist in the target function, and demonstrate in an ablation study they considerably improve performance. Our final networks are able to achieve move prediction accuracies of 41.1% and 44.4% on two different Go datasets, surpassing previous state of the art on this task by significant margins. Additionally, while previous move prediction programs have not yielded strong Go playing programs, we show that the networks trained in this work acquired high levels of skill. Our convolutional neural networks can consistently defeat the well known Go program GNU Go, indicating it is state of the art among programs that do not use Monte Carlo Tree Search. It is also able to win some games against state of the art Go playing program Fuego while using a fraction of the play time. This success at playing Go indicates high level principles of the game were learned.

cs.AI↗

Classification for Big Dataset of Bioacoustic Signals Based on Human Scoring System and Artificial Neural Network

In this paper, we propose a method to improve sound classification performance by combining signal features, derived from the time-frequency spectrogram, with human perception. The method presented herein exploits an artificial neural network (ANN) and learns the signal features based on the human perception knowledge. The proposed method is applied to a large acoustic dataset containing 24 months of nearly continuous recordings. The results show a significant improvement in performance of the detection-classification system; yielding as much as 20% improvement in true positive rate for a given false positive rate.

cs.CV↗

Bioacoustic Signal Classification Based on Continuous Region Processing, Grid Masking and Artificial Neural Network

In this paper, we develop a novel method based on machine-learning and image processing to identify North Atlantic right whale (NARW) up-calls in the presence of high levels of ambient and interfering noise. We apply a continuous region algorithm on the spectrogram to extract the regions of interest, and then use grid masking techniques to generate a small feature set that is then used in an artificial neural network classifier to identify the NARW up-calls. It is shown that the proposed technique is effective in detecting and capturing even very faint up-calls, in the presence of ambient and interfering noises. The method is evaluated on a dataset recorded in Massachusetts Bay, United States. The dataset includes 20000 sound clips for training, and 10000 sound clips for testing. The results show that the proposed technique can achieve an error rate of less than FPR = 4.5% for a 90% true positive rate.

cs.CV↗