SearcharxivSearch

arXiv subjects

Pedro Rodriguez

Publications and source records attributed to Pedro Rodriguez.

27 records · Page 2Linked to original sources

Quizbowl: The Case for Incremental Question Answering

Scholastic trivia competitions test knowledge and intelligence through mastery of question answering. Modern question answering benchmarks are one variant of the Turing test. Specifically, answering a set of questions as well as a human is a minimum bar towards demonstrating human-like intelligence. This paper makes the case that the format of one competition -- where participants can answer in the middle of hearing a question (incremental) -- better differentiates the skill between (human or machine) players. Additionally, merging a sequential decision-making sub-task with question answering (QA) provides a good setting for research in model calibration and opponent modeling. Thus, embedded in this task are three machine learning challenges: (1) factoid QA over thousands of Wikipedia-like answers, (2) calibration of the QA model's confidence scores, and (3) sequential decision-making that incorporates knowledge of the QA model, its calibration, and what the opponent may do. We make two contributions: (1) collecting and curating a large factoid QA dataset and an accompanying gameplay dataset, and (2) developing a model that addresses these three machine learning challenges. In addition to offline evaluation, we pitted our model against some of the most accomplished trivia players in the world in a series of exhibition matches spanning several years. Throughout this paper, we show that collaborations with the vibrant trivia community have contributed to the quality of our dataset, spawned new research directions, and doubled as an exciting way to engage the public with research in machine learning and natural language processing.

cs.CL

Information Seeking in the Spirit of Learning: a Dataset for Conversational Curiosity

Open-ended human learning and information-seeking are increasingly mediated by digital assistants. However, such systems often ignore the user's pre-existing knowledge. Assuming a correlation between engagement and user responses such as "liking" messages or asking followup questions, we design a Wizard-of-Oz dialog task that tests the hypothesis that engagement increases when users are presented with facts related to what they know. Through crowd-sourcing of this experiment, we collect and release 14K dialogs (181K utterances) where users and assistants converse about geographic topics like geopolitical entities and locations. This dataset is annotated with pre-existing user knowledge, message-level dialog acts, grounding to Wikipedia, and user reactions to messages. Responses using a user's prior knowledge increase engagement. We incorporate this knowledge into a multi-task model that reproduces human assistant policies and improves over a BERT content model by 13 mean reciprocal rank points.

cs.CL

Mitigating Noisy Inputs for Question Answering

Natural language processing systems are often downstream of unreliable inputs: machine translation, optical character recognition, or speech recognition. For instance, virtual assistants can only answer your questions after understanding your speech. We investigate and mitigate the effects of noise from Automatic Speech Recognition systems on two factoid Question Answering (QA) tasks. Integrating confidences into the model and forced decoding of unknown words are empirically shown to improve the accuracy of downstream neural QA systems. We create and train models on a synthetic corpus of over 500,000 noisy sentences and evaluate on two human corpora from Quizbowl and Jeopardy! competitions.

cs.CL

Data-Driven Modelling of the Van Allen Belts: The 5DRBM Model for Trapped Electrons

The magnetosphere sustained by the rotation of the Earth's liquid iron core traps charged particles, mostly electrons and protons, into structures referred to as the Van Allen belts. These radiation belts, in which the density of charged energetic particles can be very destructive for sensitive instrumentation, have to be crossed on every orbit of satellites traveling in elliptical orbits around the Earth, as is the case for ESA's INTEGRAL and XMM-Newton missions. This paper presents the first working version of the 5DRBM-e model, a global, data-driven model of the radiation belts for trapped electrons. The model is based on in-situ measurements of electrons by the radiation monitors on board the INTEGRAL and XMM-Newton satellites along their long elliptical orbits for respectively 16 and 19 years of operations. This model, in its present form, features the integral flux for trapped electrons within energies ranging from 0.7 to 1.75 MeV. Cross-validation of the 5DRBM-e with the well-known AE8min/max and AE9mean models for a low eccentricity GPS orbit shows excellent agreement, and demonstrates that the new model can be used to provide reliable predictions along widely different orbits around Earth for the purpose of designing, planning, and operating satellites with more accurate instrument safety margins. Future work will include extending the model based on electrons of different energies and proton radiation measurement data.

astro-ph.HE

Trick Me If You Can: Human-in-the-loop Generation of Adversarial Examples for Question Answering

Adversarial evaluation stress tests a model's understanding of natural language. While past approaches expose superficial patterns, the resulting adversarial examples are limited in complexity and diversity. We propose human-in-the-loop adversarial generation, where human authors are guided to break models. We aid the authors with interpretations of model predictions through an interactive user interface. We apply this generation framework to a question answering task called Quizbowl, where trivia enthusiasts craft adversarial questions. The resulting questions are validated via live human--computer matches: although the questions appear ordinary to humans, they systematically stump neural and information retrieval models. The adversarial questions cover diverse phenomena from multi-hop reasoning to entity type distractors, exposing open challenges in robust question answering.

cs.CL

Pathologies of Neural Models Make Interpretations Difficult

One way to interpret neural model predictions is to highlight the most important input features---for example, a heatmap visualization over the words in an input sentence. In existing interpretation methods for NLP, a word's importance is determined by either input perturbation---measuring the decrease in model confidence when that word is removed---or by the gradient with respect to that word. To understand the limitations of these methods, we use input reduction, which iteratively removes the least important word from the input. This exposes pathological behaviors of neural models: the remaining words appear nonsensical to humans and are not the ones determined as important by interpretation methods. As we confirm with human experiments, the reduced examples lack information to support the prediction of any label, but models still make the same predictions with high confidence. To explain these counterintuitive results, we draw connections to adversarial examples and confidence calibration: pathological behaviors reveal difficulties in interpreting neural models trained with maximum likelihood. To mitigate their deficiencies, we fine-tune the models by encouraging high entropy outputs on reduced examples. Fine-tuned models become more interpretable under input reduction without accuracy loss on regular examples.

cs.CL

Spontaneous formation of vector vortex beams in vertical-cavity surface-emitting lasers with feedback

The spontaneous emergence of vector vortex beams with non-uniform polarization distribution is reported in a vertical-cavity surface-emitting laser (VCSEL) with frequency-selective feedback. Antivortices with a hyperbolic polarization structure and radially polarized vortices are demonstrated. They exist close to and partially coexist with vortices with uniform and non-uniform polarization distributions characterized by four domains of pairwise orthogonal polarization. The spontaneous formation of these nontrivial structures in a simple, nearly isotropic VCSEL system is remarkable and the vector vortices are argued to have soliton-like properties.

physics.optics

Finding AGN in Deep X-ray Flux States with Swift

We report on our ongoing project of finding Active Galactic Nuclei (AGN) that go into deep X-ray flux states detected by Swift. Swift is performing an extensive study on the flux and spectral variability of AGN using Guest Investigator and team fill-in programs followed by triggering XMM_Newton for deeper follow-up observations. So far this program has been very successful and has led to a number of XMM-Newton follow up observations, including Mkn 335, PG 0844+349, and RX J2340.8-5329. Recent analysis of new Swift AGN observations reveal several AGN went into a very low X-ray flux state, particularly Narrow-Line Seyfert 1 galaxies. One of these is RX J2317-4422, which dropped by a factor of about 60 when compared to the ROSAT All-Sky Survey.

astro-ph.HE

Solar Control on Jupiter's Equatorial X-ray Emissions: 26-29 November 2003 XMM-Newton Observation

During November 26-29, 2003 XMM-Newton observed soft (0.2-2 keV) X-ray emission from Jupiter for 69 hours. The low-latitude X-ray disk emission of Jupiter is observed to be almost uniform in intensity with brightness that is consistent with a solar-photon driven process. The simultaneous lightcurves of Jovian equatorial X-rays and solar X-rays (measured by the TIMED/SEE and GOES satellites) show similar day-to-day variability. A large solar X-ray flare occurring on the Jupiter-facing side of the Sun is found to have a corresponding feature in the Jovian X-rays. These results support the hypothesis that X-ray emission from Jovian low-latitudes are solar X-rays scattered from the planet's upper atmosphere, and suggest that the Sun directly controls the non-auroral X-rays from Jupiter's disk. Our study also suggests that Jovian equatorial X-rays can be used to monitor the solar X-ray flare activity on the hemisphere of the Sun that is invisible to space weather satellites.

astro-ph