Searcharxiv⌕ Search

arXiv subjects

Howard Chen

Publications and source records attributed to Howard Chen.

At least 37 records · Page 2Linked to original sources

What In-Context Learning "Learns" In-Context: Disentangling Task Recognition and Task Learning

Large language models (LLMs) exploit in-context learning (ICL) to solve tasks with only a few demonstrations, but its mechanisms are not yet well-understood. Some works suggest that LLMs only recall already learned concepts from pre-training, while others hint that ICL performs implicit learning over demonstrations. We characterize two ways through which ICL leverages demonstrations. Task recognition (TR) captures the extent to which LLMs can recognize a task through demonstrations -- even without ground-truth labels -- and apply their pre-trained priors, whereas task learning (TL) is the ability to capture new input-label mappings unseen in pre-training. Using a wide range of classification datasets and three LLM families (GPT-3, LLaMA and OPT), we design controlled experiments to disentangle the roles of TR and TL in ICL. We show that (1) models can achieve non-trivial performance with only TR, and TR does not further improve with larger models or more demonstrations; (2) LLMs acquire TL as the model scales, and TL's performance consistently improves with more demonstrations in context. Our findings unravel two different forces behind ICL and we advocate for discriminating them in future ICL research due to their distinct nature.

cs.CL↗

Sporadic Spin-Orbit Variations in Compact Multi-planet Systems and their Influence on Exoplanet Climate

Climate modeling has shown that tidally influenced terrestrial exoplanets, particularly those orbiting M-dwarfs, have unique atmospheric dynamics and surface conditions that may enhance their likelihood to host viable habitats. However, sporadic libration and rotation induced by planetary interactions, such as that due to mean motion resonances (MMRs) in compact planetary systems may destabilize attendant exoplanets away from synchronized states (or 1:1 spin-orbit ratio). Here, we use a three-dimensional N-Rigid-Body integrator and an intermediately-complex general circulation model to simulate the evolving climates of TRAPPIST-1 e and f with different orbital and spin evolution pathways. Planet f perturbed by MMR effects with chaotic spin-variations are colder and dryer compared to their synchronized counterparts due to the zonal drift of the substellar point away from open ocean basins of their initial eyeball states. On the other hand, the differences between perturbed and synchronized planet e are minor due to higher instellation, warmer surfaces, and reduced climate hysteresis. This is the first study to incorporate the time-dependent outcomes of direct gravitational N-Rigid-Body simulations into 3D climate modeling of extrasolar planets and our results show that planets at the outer edge of the habitable zones in compact multiplanet systems are vulnerable to rapid global glaciations. In the absence of external mechanisms such as orbital forcing or tidal heating, these planets could be trapped in permanent snowball states.

astro-ph.EP↗

WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Existing benchmarks for grounding language in interactive environments either lack real-world linguistic elements, or prove difficult to scale up due to substantial human involvement in the collection of data or feedback signals. To bridge this gap, we develop WebShop -- a simulated e-commerce website environment with $1.18$ million real-world products and $12,087$ crowd-sourced text instructions. Given a text instruction specifying a product requirement, an agent needs to navigate multiple types of webpages and issue diverse actions to find, customize, and purchase an item. WebShop provides several challenges for language grounding including understanding compositional instructions, query (re-)formulation, comprehending and acting on noisy text in webpages, and performing strategic exploration. We collect over $1,600$ human demonstrations for the task, and train and evaluate a diverse range of agents using reinforcement learning, imitation learning, and pre-trained image and language models. Our best model achieves a task success rate of $29\%$, which outperforms rule-based heuristics ($9.6\%$) but is far lower than human expert performance ($59\%$). We also analyze agent and human trajectories and ablate various model components to provide insights for developing future agents with stronger language understanding and decision making abilities. Finally, we show that agents trained on WebShop exhibit non-trivial sim-to-real transfer when evaluated on amazon.com and ebay.com, indicating the potential value of WebShop in developing practical web-based agents that can operate in the wild.

cs.CL↗

Controllable Text Generation with Language Constraints

We consider the task of text generation in language models with constraints specified in natural language. To this end, we first create a challenging benchmark Cognac that provides as input to the model a topic with example text, along with a constraint on text to be avoided. Unlike prior work, our benchmark contains knowledge-intensive constraints sourced from databases like Wordnet and Wikidata, which allows for straightforward evaluation while striking a balance between broad attribute-level and narrow lexical-level controls. We find that even state-of-the-art language models like GPT-3 fail often on this task, and propose a solution to leverage a language model's own internal knowledge to guide generation. Our method, called CognacGen, first queries the language model to generate guidance terms for a specified topic or constraint, and uses the guidance to modify the model's token generation probabilities. We propose three forms of guidance (binary verifier, top-k tokens, textual example), and employ prefix-tuning approaches to distill the guidance to tackle diverse natural language constraints. Through extensive empirical evaluations, we demonstrate that CognacGen can successfully generalize to unseen instructions and outperform competitive baselines in generating constraint conforming text.

cs.CL↗

Impact Induced Atmosphere-Mantle Exchange Sets the Volatile Elemental Ratios on Primitive Earths

Conventional planet formation theory suggests that chondritic materials have delivered crucial atmospheric and hydrospheric elements such as carbon (C), nitrogen (N), and hydrogen (H) onto primitive Earth. However, recent measurements highlight the significant elemental ratio discrepancies between terrestrial parent bodies and the supposed planet building blocks. Here we present a volatile evolution model during the assembly of Earth and Earth-like planets. Our model includes impact losses, atmosphere-mantle exchange, and time dependent effects of accretion and outgassing calculated from dynamical modeling outcomes. Exploring a wide range of planetesimal properties (i.e., size and composition) as well as impact history informed by N-body accretion simulations, we find that while the degree of CNH fractionation has inherent stochasticity, the evolution of C/N and C/H ratios can be traced back to certain properties of the protoplanet and projectiles. Interestingly, the majority of our Earth-like planets acquire superchondritic final C/N ratios, implying that the volatile elemental ratios on terrestrial planets are driven by the complex interplay between delivery, atmospheric ablation, and mantle degassing.

astro-ph.EP↗

Can Rationalization Improve Robustness?

A growing line of work has investigated the development of neural NLP models that can produce rationales--subsets of input that can explain their model predictions. In this paper, we ask whether such rationale models can also provide robustness to adversarial attacks in addition to their interpretable nature. Since these models need to first generate rationales ("rationalizer") before making predictions ("predictor"), they have the potential to ignore noise or adversarially added text by simply masking it out of the generated rationale. To this end, we systematically generate various types of 'AddText' attacks for both token and sentence-level rationalization tasks, and perform an extensive empirical evaluation of state-of-the-art rationale models across five different tasks. Our experiments reveal that the rationale models show the promise to improve robustness, while they struggle in certain scenarios--when the rationalizer is sensitive to positional bias or lexical choices of attack text. Further, leveraging human rationale as supervision does not always translate to better performance. Our study is a first step towards exploring the interplay between interpretability and robustness in the rationalize-then-predict framework.

cs.CL↗

Habitability Models for Astrobiology

Habitability has been generally defined as the capability of an environment to support life. Ecologists have been using Habitat Suitability Models (HSMs) for more than four decades to study the habitability of Earth from local to global scales. Astrobiologists have been proposing different habitability models for some time, with little integration and consistency among them, being different in function to those used by ecologists. Habitability models are not only used to determine if environments are habitable or not, but they also are used to characterize what key factors are responsible for the gradual transition from low to high habitability states. Here we review and compare some of the different models used by ecologists and astrobiologists and suggest how they could be integrated into new habitability standards. Such standards will help to improve the comparison and characterization of potentially habitable environments, prioritize target selections, and study correlations between habitability and biosignatures. Habitability models are the foundation of planetary habitability science and the synergy between ecologists and astrobiologists is necessary to expand our understanding of the habitability of Earth, the Solar System, and extrasolar planets.

astro-ph.EP↗

Persistence of Flare-Driven Atmospheric Chemistry on Rocky Habitable Zone Worlds

Low-mass stars show evidence of vigorous magnetic activity in the form of large flares and coronal mass ejections. Such space weather events may have important ramifications for the habitability and observational fingerprints of exoplanetary atmospheres. Here, using a suite of three-dimensional coupled chemistry-climate model (CCM) simulations, we explore effects of time-dependent stellar activity on rocky planet atmospheres orbiting G-, K-, and M-dwarf stars. We employ observed data from the MUSCLES campaign and Transiting Exoplanet Satellite Survey and test a range of rotation period, magnetic field strength, and flare frequency assumptions. We find that recurring flares drive K- and M-dwarf planet atmospheres into chemical equilibria that substantially deviate from their pre-flare regimes, whereas G-dwarf planet atmospheres quickly return to their baseline states. Interestingly, simulated O$_2$-poor and O$_2$-rich atmospheres experiencing flares produce similar mesospheric nitric oxide abundances, suggesting that stellar flares can highlight otherwise undetectable chemical species. Applying a radiative transfer model to our CCM results, we find that flare-driven transmission features of bio-indicating species, such as nitrogen dioxide, nitrous oxide, and nitric acid, show particular promise for detection by future instruments.

astro-ph.EP↗

Non-Parametric Few-Shot Learning for Word Sense Disambiguation

Word sense disambiguation (WSD) is a long-standing problem in natural language processing. One significant challenge in supervised all-words WSD is to classify among senses for a majority of words that lie in the long-tail distribution. For instance, 84% of the annotated words have less than 10 examples in the SemCor training data. This issue is more pronounced as the imbalance occurs in both word and sense distributions. In this work, we propose MetricWSD, a non-parametric few-shot learning approach to mitigate this data imbalance issue. By learning to compute distances among the senses of a given word through episodic training, MetricWSD transfers knowledge (a learned metric space) from high-frequency words to infrequent ones. MetricWSD constructs the training episodes tailored to word frequencies and explicitly addresses the problem of the skewed distribution, as opposed to mixing all the words trained with parametric models in previous work. Without resorting to any lexical resources, MetricWSD obtains strong performance against parametric alternatives, achieving a 75.1 F1 score on the unified WSD evaluation benchmark (Raganato et al., 2017b). Our analysis further validates that infrequent words and senses enjoy significant improvement.

cs.CL↗

TRAPPIST Habitable Atmosphere Intercomparison (THAI) workshop report

The era of atmospheric characterization of terrestrial exoplanets is just around the corner. Modeling prior to observations is crucial in order to predict the observational challenges and to prepare for the data interpretation. This paper presents the report of the TRAPPIST Habitable Atmosphere Intercomparison (THAI) workshop (14-16 September 2020). A review of the climate models and parameterizations of the atmospheric processes on terrestrial exoplanets, model advancements and limitations, as well as direction for future model development was discussed. We hope that this report will be used as a roadmap for future numerical simulations of exoplanet atmospheres and maintaining strong connections to the astronomical community.

astro-ph.EP↗

Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems

Existing goal-oriented dialogue datasets focus mainly on identifying slots and values. However, customer support interactions in reality often involve agents following multi-step procedures derived from explicitly-defined company policies as well. To study customer service dialogue systems in more realistic settings, we introduce the Action-Based Conversations Dataset (ABCD), a fully-labeled dataset with over 10K human-to-human dialogues containing 55 distinct user intents requiring unique sequences of actions constrained by policies to achieve task success. We propose two additional dialog tasks, Action State Tracking and Cascading Dialogue Success, and establish a series of baselines involving large-scale, pre-trained language models on this dataset. Empirical results demonstrate that while more sophisticated networks outperform simpler models, a considerable gap (50.8% absolute accuracy) still exists to reach human-level performance on ABCD.

cs.CL↗

Autoregressive Knowledge Distillation through Imitation Learning

The performance of autoregressive models on natural language generation tasks has dramatically improved due to the adoption of deep, self-attentive architectures. However, these gains have come at the cost of hindering inference speed, making state-of-the-art models cumbersome to deploy in real-world, time-sensitive settings. We develop a compression technique for autoregressive models that is driven by an imitation learning perspective on knowledge distillation. The algorithm is designed to address the exposure bias problem. On prototypical language generation tasks such as translation and summarization, our method consistently outperforms other distillation algorithms, such as sequence-level knowledge distillation. Student models trained with our method attain 1.4 to 4.8 BLEU/ROUGE points higher than those trained from scratch, while increasing inference speed by up to 14 times in comparison to the teacher model.

cs.CL↗

Habitability Models for Planetary Sciences

Habitability has been generally defined as the capability of an environment to support life. Ecologists have been using Habitat Suitability Models (HSMs) for more than four decades to study the habitability of Earth from local to global scales. Astrobiologists have been proposing different habitability models for some time, with little integration and consistency between them and different in function to those used by ecologists. In this white paper, we suggest a mass-energy habitability model as an example of how to adapt and expand the models used by ecologists to the astrobiology field. We propose to implement these models into a NASA Habitability Standard (NHS) to standardize the habitability objectives of planetary missions. These standards will help to compare and characterize potentially habitable environments, prioritize target selections, and study correlations between habitability and biosignatures. Habitability models are the foundation of planetary habitability science. The synergy between the methods used by ecologists and astrobiologists will help to integrate and expand our understanding of the habitability of Earth, the Solar System, and exoplanets.

astro-ph.IM↗

Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments

We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task. We introduce the Touchdown task and dataset, where an agent must first follow navigation instructions in a real-life visual urban environment, and then identify a location described in natural language to find a hidden object at the goal position. The data contains 9,326 examples of English instructions and spatial descriptions paired with demonstrations. Empirical analysis shows the data presents an open challenge to existing methods, and qualitative linguistic analysis shows that the data displays richer use of spatial reasoning compared to related resources.

cs.CV↗

Interactive Classification by Asking Informative Questions

We study the potential for interaction in natural language classification. We add a limited form of interaction for intent classification, where users provide an initial query using natural language, and the system asks for additional information using binary or multi-choice questions. At each turn, our system decides between asking the most informative question or making the final classification prediction.The simplicity of the model allows for bootstrapping of the system without interaction data, instead relying on simple crowdsourcing tasks. We evaluate our approach on two domains, showing the benefit of interaction and the advantage of learning to balance between asking additional questions and making the final prediction.

cs.CL↗

Habitability and Spectroscopic Observability of Warm M-dwarf Exoplanets Evaluated with a 3D Chemistry-Climate Model

Planets residing in circumstellar habitable zones (CHZs) offer our best opportunities to test hypotheses of life's potential pervasiveness and complexity. Constraining the precise boundaries of habitability and its observational discriminants is critical to maximizing our chances at remote life detection with future instruments. Conventionally, calculations of the inner edge of the habitable zone (IHZ) have been performed using both 1D radiative-convective and 3D general circulation models. However, these models lack interactive three-dimensional chemistry and do not resolve the mesosphere and lower thermosphere (MLT) region of the upper atmosphere. Here we employ a 3D high-top chemistry-climate model (CCM) to simulate the atmospheres of synchronously-rotating planets orbiting at the inner edge of habitable zones of K- and M-dwarf stars (between $T_{\rm eff} =$ 2600 K and 4000 K). While our IHZ climate predictions are in good agreement with GCM studies, we find noteworthy departures in simulated ozone and HO$_{\rm x}$ photochemistry. For instance, climates around inactive stars do not typically enter the classical moist greenhouse regime even with high ($< 10^{-3}$ mol mol$^{-1}$) stratospheric water vapor mixing ratios, which suggests that planets around inactive M-stars may only experience minor water-loss over geologically significant timescales. In addition, we find much thinner ozone layers on potentially habitable moist greenhouse atmospheres, as ozone experiences rapid destruction via reaction with hydrogen oxide radicals. Using our CCM results as inputs, our simulated transmission spectra show that both water vapor and ozone features on moist greenhouse atmospheres could be detectable by instruments NIRSpec and MIRI LRS onboard the James Webb Space Telescope.

astro-ph.EP↗

Biosignature Anisotropy Modeled on Temperate Tidally Locked M-dwarf Planets

A planet's atmospheric constituents (e.g., O$_2$, O$_3$, H$_2$O$_v$, CO$_2$, CH$_4$, N$_2$O) can provide clues to its surface habitability, and may offer biosignature targets for remote life detection efforts. The plethora of rocky exoplanets found by recent transit surveys (e.g., the Kepler mission) indicates that potentially habitable systems orbiting K- and M-dwarf stars may have very different orbital and atmospheric characteristics than Earth. To assess the physical distribution and observational prospects of various biosignatures and habitability indicators, it is important to understand how they may change under different astrophysical and geophysical configurations, and to simulate these changes with models that include feedbacks between different subsystems of a planet's climate. Here we use a three-dimensional (3D) Chemistry-Climate model (CCM) to study the effects of changes in stellar spectral energy distribution (SED), stellar activity, and planetary rotation on Earth-analogs and tidally-locked planets. Our simulations show that, apart from shifts in stellar SEDs and UV radiation, changes in illumination geometry and rotation-induced circulation can influence the global distribution of atmospheric biosignatures. We find that the stratospheric day-to-night side mixing ratio differences on tidally-locked planets remain low ($<20\%$) across the majority of the canonical biosignatures. Interestingly however, secondary photosynthetic biosignatures (e.g., C$_2$H$_6$S) show much greater (${\sim}67\%$) day-to-night side differences, and point to regimes in which tidal-locking could have observationally distinguishable effects on phase curve, transit, and secondary eclipse measurements. Overall, this work highlights the potential and promise for 3D CCMs to study the atmospheric properties and habitability of terrestrial worlds.

astro-ph.EP↗

Habitable Evaporated Cores and the Occurrence of Panspermia near the Galactic Center

Black holes growing via the accretion of gas emit radiation that can photoevaporate the atmospheres of nearby planets. Here we couple planetary structural evolution models of sub-Neptune mass planets to the growth of the Milky way's central supermassive black-hole, Sgr A$^*$ and investigate how planetary evolution is influenced by quasar activity. We find that, out to ${\sim} 20$ pc from Sgr A$^*$, the XUV flux emitted during its quasar phase can remove several percent of a planet's H/He envelope by mass; in many cases, this removal results in bare rocky cores, many of which situated in the habitable zones (HZs) of G-type stars. The erosion of sub-Neptune sized planets may be one of the most prevalent channels by which terrestrial super-Earths are created near the Galactic Center. As such, the planet population demographics may be quite different close to Sgr A$^*$ than in the Galaxy's outskirts. The high stellar densities in this region (about seven orders of magnitude greater than the solar neighborhood) imply that the distance between neighboring rocky worlds is short ($500-5000$~AU). The proximity between potentially habitable terrestrial planets may enable the onset of widespread interstellar panspermia near the nuclei of galaxies. More generally, we predict these phenomena to be ubiquitous for planets in nuclear star clusters and ultra-compact dwarfs. Globular clusters, on the other hand, are less affected by the black holes.

astro-ph.EP↗