SearcharxivSearch

arXiv subjects

Howard Chen

Publications and source records attributed to Howard Chen.

At least 19 recordsLinked to original sources

Multiple Double Arithmetic on NVIDIA Tensor Cores

A multiple double is an unevaluated sum of doubles. An NVIDIA tensor core is a specialized high performance compute core for matrix multiplication. The Ampere A100, released in 2020, introduced tensor cores capable of 64-bit floating-point arithmetic. Every multiple double arithmetical operation requires renormalization, which involves branching, for which tensor cores are unsuited. To solve this problem caused by renormalization, we apply a solution similar to the Ozaki scheme [Ozaki et al, Numerical Algorithms, 2012]. Our software is available under the GPU GPL license on github.

cs.MS

An Adolescent and Near-Resonant Planetary System Near the End of Photoevaporation

Young exoplanets provide vital insights into the early dynamical and atmospheric evolution of planetary systems. Many multi-planet systems younger than 100 Myr exhibit mean-motion resonances, likely established through convergent disk migration. Over time, however, these resonant chains are often disrupted, mirroring the Nice model proposed for the Solar System. We present a detailed characterization of the ~200-Myr-old TOI-2076 system, which contains four sub-Neptune planets between 1.4 and 3.5 Earth radii. We demonstrate that its planets are near but not locked in mean-motion resonances, making the system dynamically fragile. The four planets have comparable core masses but display a monotonic increase in hydrogen and helium (H/He) envelope mass fractions (stripped-1%-5%-5%) with decreasing stellar insolation. This trend is consistent with atmospheric mass-loss due to photoevaporation, which predicts that the envelopes of irradiated planets either erode completely or stabilize at a residual level of ~1% by mass within the first few hundred million years, with more distant, less-irradiated planets retaining most of primordial envelopes. Additionally, previous detections of metastable helium outflows rule out a pure water-world scenario for TOI-2076 planets. Our finding provides direct observational evidence that the dynamical and atmospheric reshaping of compact planetary systems begin early, offering an empirical anchor for models of their long-term evolution.

astro-ph.EP

Topography-Induced Stationary Waves and the Onset of Nightside Warming on Rocky Planets around M-dwarf Stars

Among potentially habitable worlds, rocky planets orbiting M dwarfs offer the most favorable prospects for atmospheric characterization, yet their climates may differ substantially from those of Earth analogs. In the tidally locked limit, the nightside's tendency to radiatively cool and potentially trap volatiles as permanent ice introduces a strong dependence of habitability on the planet's surface and atmospheric boundary conditions. We perform a suite of synchronously rotating experiments spanning a wide range of topographic and orographic realizations with different mean elevations and landmass distributions. Across a grid of $p_{\mathrm{N2}} = 0.5$-$8~\mathrm{bar}$ and $F_{\star} = 1200$-$1700~\mathrm{W\,m^{-2}}$, we find that surface relief breaks the flow symmetry, replacing the circumpolar vortices with mechanically forced stationary waves. Steep orography produces standing Rossby gyres that strengthen the cross-terminator jet and align vertical uplift with the day--night boundary. These new circulation regimes enhance moisture transport, increasing the infrared optical depth and promoting additional nightside cloud formation, which produces a stronger cloud-greenhouse feedback and lower the critical fluxes required for global planetary deglaciation. Broad, elevated plateaus drive a similarly fragmented but slightly weaker circulation, yielding less effective moisture transport. These results show that the relief and spatial distribution of landmasses, parameters unconstrained for most exoplanets, can exert strong controls on the climatic bifurcations of tidally locked M-dwarf exoplanets.

astro-ph.EP

What's Inside Matters: The Effect of Oxygen Fugacity and Initial Volatile Abundance on the Atmospheres of the TRAPPIST-1 Planets

The TRAPPIST-1 planets have become prime targets for studying the atmospheric and geophysical properties of planets around M-dwarf stars. To effectively identify their atmospheric composition, we first must understand their geological evolution. For this study, we focus on enhancing an existing atmosphere-interior exchange model by incorporating additional geological processes relevant to rocky planets. We have extended the model to include the carbon cycle, which enables the model to track four key gas species - CO$_2$, CO, H$_2$O, and H$_2$ - across four planetary reservoirs: the mantle, plate, ocean, and atmosphere. Major features added include surface temperature calculations which are crucial for the carbon cycle, oxygen fugacity as a planetary interior parameter in the model, and oxidation reactions and diffusion-limited escape calculations to the atmosphere portion of the model. We successfully validated the model for Earth and applied this model to study the effect of oxygen fugacity and initial water abundance on TRAPPIST-1 d, e and f. Our results for present-day abundances show that oxygen fugacity significantly affects the partial pressures of H$_2$ and CO$_2$ for all three planets with minor effects for CO on two of the planets. We also found that H$_2$ is strongly dependent on water mass fraction (WMF). The addition of atmospheric processes produced a significant difference in the H$_2$ and CO abundances at present-day. These results highlight the importance of considering interior parameters to be able to further constrain the geological evolution of these planets and effectively put atmosphere observations into context.

astro-ph.EP

Accumulating Context Changes the Beliefs of Language Models

Language model (LM) assistants are increasingly used in applications such as brainstorming and research. Improvements in memory and context size have allowed these models to become more autonomous, which has also resulted in more text accumulation in their context windows without explicit user intervention. This comes with a latent risk: the belief profiles of models -- their understanding of the world as manifested in their responses or actions -- may silently change as context accumulates. This can lead to subtly inconsistent user experiences, or shifts in behavior that deviate from the original alignment of the models. In this paper, we explore how accumulating context by engaging in interactions and processing text -- talking and reading -- can change the beliefs of language models, as manifested in their responses and behaviors. Our results reveal that models' belief profiles are highly malleable: GPT-5 exhibits a 54.7% shift in its stated beliefs after 10 rounds of discussion about moral dilemmas and queries about safety, while Grok 4 shows a 27.2% shift on political issues after reading texts from the opposing position. We also examine models' behavioral changes by designing tasks that require tool use, where each tool selection corresponds to an implicit belief. We find that these changes align with stated belief shifts, suggesting that belief shifts will be reflected in actual behavior in agentic systems. Our analysis exposes the hidden risk of belief shift as models undergo extended sessions of talking or reading, rendering their opinions and actions unreliable.

cs.CL

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting

Adapting language models (LMs) to new tasks via post-training carries the risk of degrading existing capabilities -- a phenomenon classically known as catastrophic forgetting. In this paper, toward identifying guidelines for mitigating this phenomenon, we systematically compare the forgetting patterns of two widely adopted post-training methods: supervised fine-tuning (SFT) and reinforcement learning (RL). Our experiments reveal a consistent trend across LM families (Llama, Qwen) and tasks (instruction following, general knowledge, and arithmetic reasoning): RL leads to less forgetting than SFT while achieving comparable or higher target task performance. To investigate the cause for this difference, we consider a simplified setting in which the LM is modeled as a mixture of two distributions, one corresponding to prior knowledge and the other to the target task. We identify that the mode-seeking nature of RL, which stems from its use of on-policy data, enables keeping prior knowledge intact when learning the target task. We then verify this insight by demonstrating that the use on-policy data underlies the robustness of RL to forgetting in practical settings, as opposed to other algorithmic choices such as the KL regularization or advantage estimation. Lastly, as a practical implication, our results highlight the potential of mitigating forgetting using approximately on-policy data, which can be substantially more efficient to obtain than fully on-policy data.

cs.LG

Born Dry or Born Wet? A Palette of Water Growth Histories in TRAPPIST-1 Analogs and Compact Planetary Systems

It is still unclear whether exoplanets in compact multiplanet systems such as TRAPPIST-1 are able to accrete large quantities of volatiles, grow to sufficient mass, and maintain robust atmospheres and hydrospheres. Previous estimates of water content in M-dwarf systems have largely relied on population synthesis or atmosphere-interior evolution models, often treating impacts and atmospheric loss in isolation. In this work, we couple impact delivery, impact erosion, and mantle-atmosphere exchange within a model that tracks volatile evolution through stochastic collision histories. By explicitly including both planetesimal accretion and the prolonged luminous pre-main-sequence phase of M dwarfs, we find lower water inventories for the inner TRAPPIST-1 analogs (b-e), spanning only $10^{-4}$-$10^{-2} M_{\oplus,\rm ocn}$ across a wide range of disk structures and impact scenarios. By contrast, the outer planets (f-h analogs) frequently retain water inventories exceeding an Earth ocean mass. This systematic volatile gradient provides a physically motivated explanation for JWST's nondetections of atmospheres on TRAPPIST-1 b and c, implying an origin rooted in formation conditions rather than in post-formation escape. Our results suggest that many rocky planets in compact M-dwarf systems may form already depleted in volatile compounds, fundamentally limiting their capacity to sustain atmospheres or surface oceans. More broadly, our multistage framework for volatile tracking can help interpret future observations of compact systems and set more realistic initial conditions for exoplanet interior compositions and atmospheric models.

astro-ph.EP

Statutory Construction and Interpretation for Artificial Intelligence

AI systems are increasingly governed by natural language principles, yet a key challenge arising from reliance on language remains underexplored: interpretive ambiguity. As in legal systems, ambiguity arises both from how these principles are written and how they are applied. But while legal systems use institutional safeguards to manage such ambiguity, such as transparent appellate review policing interpretive constraints, AI alignment pipelines offer no comparable protections. Different interpretations of the same rule can lead to inconsistent or unstable model behavior. Drawing on legal theory, we identify key gaps in current alignment pipelines by examining how legal systems constrain ambiguity at both the rule creation and rule application steps. We then propose a computational framework that mirrors two legal mechanisms: (1) a rule refinement pipeline that minimizes interpretive disagreement by revising ambiguous rules (analogous to agency rulemaking or iterative legislative action), and (2) prompt-based interpretive constraints that reduce inconsistency in rule application (analogous to legal canons that guide judicial discretion). We evaluate our framework on a 5,000-scenario subset of the WildChat dataset and show that both interventions significantly improve judgment consistency across a panel of reasonable interpreters. Our approach offers a first step toward systematically managing interpretive ambiguity, an essential step for building more robust, law-following AI systems.

cs.CL

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Using AI to create autonomous researchers has the potential to accelerate scientific discovery. A prerequisite for this vision is understanding how well an AI model can identify the underlying structure of a black-box system from its behavior. In this paper, we explore how well a large language model (LLM) learns to identify a black-box function from passively observed versus actively collected data. We investigate the reverse-engineering capabilities of LLMs across three distinct types of black-box systems, each chosen to represent different problem domains where future autonomous AI researchers may have considerable impact: Program, Formal Language, and Math Equation. Through extensive experiments, we show that LLMs fail to extract information from observations, reaching a performance plateau that falls short of the ideal of Bayesian inference. However, we demonstrate that prompting LLMs to not only observe but also intervene -- actively querying the black-box with specific inputs to observe the resulting output -- improves performance by allowing LLMs to test edge cases and refine their beliefs. By providing the intervention data from one LLM to another, we show that this improvement is partly a result of engaging in the process of generating effective interventions, paralleling results in the literature on human learning. Further analysis reveals that engaging in intervention can help LLMs escape from two common failure modes: overcomplication, where the LLM falsely assumes prior knowledge about the black-box, and overlooking, where the LLM fails to incorporate observations. These insights provide practical guidance for helping LLMs more effectively reverse-engineer black-box systems, supporting their use in making new discoveries.

cs.LG

Effects of transient stellar emissions on planetary climates of tidally-locked exo-earths

Space weather in exoplanetary systems, driven by transient stellar emissions such as flares, coronal mass ejections, and stellar proton events, can significantly influence planetary habitability and the long-term evolution of atmospheres. These time-dependent phenomena also complicate the remote characterization of exoplanets by altering the abundance of key chemical species and modulating atmospheric brightness temperatures. While prior studies have largely focused on photochemical effects, surface UV dosages, and spectral consequences, here we extend the analysis using three-dimensional general circulation models coupled with interactive photochemistry. We simulate the climate and chemical responses of TRAPPIST-1e-like, synchronously rotating planets subjected to stellar energetic particle events and periodic UV flux enhancements. Using statistical methods, we evaluate impacts across spatial and temporal scales. Our results show that abrupt thermospheric cooling occurs via radiative emissions from NO and CO2, while warming in the middle and lower atmosphere arises from increased infrared absorbers, including N2O and H2O. In moderately active stellar regimes, atmospheric temperature changes are strongly modulated by O3 variability. Cumulative effects depend on flare frequency, while instantaneous responses are sensitive to the spectral energy distribution of the flare. Notably, intense flares can dynamically energize the middle atmosphere, enhancing wind speeds by up to 40 m/s on the substellar nightside at altitudes of 30 to 50 km. These findings suggest that repeated, high-energy eruptive events from young stars may play a critical role in shaping atmospheric dynamics on temperate terrestrial exoplanets.

astro-ph.EP

Continual Memorization of Factoids in Language Models

As new knowledge rapidly accumulates, language models (LMs) with pretrained knowledge quickly become obsolete. A common approach to updating LMs is fine-tuning them directly on new knowledge. However, recent studies have shown that fine-tuning for memorization may be ineffective in storing knowledge or may exacerbate hallucinations. In this work, we introduce a setting we call continual memorization, where a model must memorize and retain a set of factoids through multiple stages of fine-tuning on subsequent datasets. We characterized the forgetting patterns through extensive experiments and show that LMs widely suffer from forgetting, especially when needing to memorize factoids in the second stage. We posit that forgetting can be alleviated by modifying training dynamics: (1) protecting the memorization process when learning factoids or (2) reducing interference from subsequent training stages. Intriguingly, we find that mixing randomly generated word sequences or generic data sampled from pretraining corpora at different training stages effectively mitigates forgetting REMIX: Random and Generic Data Mixing). REMIX can recover performance from severe forgetting, outperforming replay methods and other continual learning baselines. We analyze how REMIX influences the learning process and find that robust memorization follows a distinct pattern: the model stores factoids in earlier layers than usual and diversifies the layers that retain them, which results in easier recall and manipulate of the learned factoids.

cs.CL

CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an over-optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions can deteriorate performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project page and leaderboard: https://charxiv.github.io/

cs.CL

Language Models as Science Tutors

NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations.

cs.CL

Walking Down the Memory Maze: Beyond Context Limit through Interactive Reading

Large language models (LLMs) have advanced in large strides due to the effectiveness of the self-attention mechanism that processes and compares all tokens at once. However, this mechanism comes with a fundamental issue -- the predetermined context window is bound to be limited. Despite attempts to extend the context window through methods like extrapolating the positional embedding, using recurrence, or selectively retrieving essential parts of the long sequence, long-text understanding continues to be a challenge. We propose an alternative approach which instead treats the LLM as an interactive agent, allowing it to decide how to read the text via iterative prompting. We introduce MemWalker, a method that first processes the long context into a tree of summary nodes. Upon receiving a query, the model navigates this tree in search of relevant information, and responds once it gathers sufficient information. On long-text question answering tasks our method outperforms baseline approaches that use long context windows, recurrence, and retrieval. We show that, beyond effective reading, MemWalker enhances explainability by highlighting the reasoning steps as it interactively reads the text; pinpointing the relevant text segments related to the query.

cs.CL

Deuterium Escape on Photoevaporating Sub-Neptunes

We investigate the evolution of the deuterium-to-hydrogen (D/H) mass ratio driven by EUV photoevaporation of hydrogen-rich atmospheres of close-in sub-Neptunes around solar-type stars. For the first time, the diffusion-limited approach in conjunction with energy-limited photoevaporation is considered in evaluating deuterium escape from evolving exoplanet H/He envelopes. We find that the planets with smaller initial gas envelopes and thus smaller sizes can lead to weaker atmospheric escape, which facilitates hydrogen-deuterium fractionation. Specifically, in our grid of simulations with a low envelope mass fraction less than 0.005, a low-mass sub-Neptune (4-$5M_\oplus$) at about 0.25-0.4 au or a high-mass sub-Neptune (10-$15M_\oplus$) at about 0.1-0.25 au can increase the D/H values by greater than 20% over 7.5 Gyr. Akin to the helium-enhanced envelopes of sub-Neptunes due to photoevaporating escape, the planets along the upper boundary of the radius valley are the best targets to detect high D/H ratios. The ratio can rise by a factor of $\lesssim$ 1.65 within 7.5 Gyrs in our grid of evolutionary calculations. The D/H ratio is expected to be higher in thinner envelopes as long as the planets do not become bare rocky cores.

astro-ph.EP

COLLIE: Systematic Construction of Constrained Text Generation Tasks

Text generation under constraints have seen increasing interests in natural language processing, especially with the rapidly improving capabilities of large language models. However, existing benchmarks for constrained generation usually focus on fixed constraint types (e.g.,generate a sentence containing certain words) that have proved to be easy for state-of-the-art models like GPT-4. We present COLLIE, a grammar-based framework that allows the specification of rich, compositional constraints with diverse generation levels (word, sentence, paragraph, passage) and modeling challenges (e.g.,language understanding, logical reasoning, counting, semantic planning). We also develop tools for automatic extraction of task instances given a constraint structure and a raw text corpus. Using COLLIE, we compile the COLLIE-v1 dataset with 2080 instances comprising 13 constraint structures. We perform systematic experiments across five state-of-the-art instruction-tuned language models and analyze their performances to reveal shortcomings. COLLIE is designed to be extensible and lightweight, and we hope the community finds it useful to develop more complex constraints and evaluations in the future.

cs.CL

C-STS: Conditional Semantic Textual Similarity

Semantic textual similarity (STS), a cornerstone task in NLP, measures the degree of similarity between a pair of sentences, and has broad application in fields such as information retrieval and natural language understanding. However, sentence similarity can be inherently ambiguous, depending on the specific aspect of interest. We resolve this ambiguity by proposing a novel task called Conditional STS (C-STS) which measures sentences' similarity conditioned on an feature described in natural language (hereon, condition). As an example, the similarity between the sentences "The NBA player shoots a three-pointer." and "A man throws a tennis ball into the air to serve." is higher for the condition "The motion of the ball" (both upward) and lower for "The size of the ball" (one large and one small). C-STS's advantages are two-fold: (1) it reduces the subjectivity and ambiguity of STS and (2) enables fine-grained language model evaluation through diverse natural language conditions. We put several state-of-the-art models to the test, and even those performing well on STS (e.g. SimCSE, Flan-T5, and GPT-4) find C-STS challenging; all with Spearman correlation scores below 50. To encourage a more comprehensive evaluation of semantic similarity and natural language understanding, we make nearly 19K C-STS examples and code available for others to train and test their models.

cs.CL

What In-Context Learning "Learns" In-Context: Disentangling Task Recognition and Task Learning

Large language models (LLMs) exploit in-context learning (ICL) to solve tasks with only a few demonstrations, but its mechanisms are not yet well-understood. Some works suggest that LLMs only recall already learned concepts from pre-training, while others hint that ICL performs implicit learning over demonstrations. We characterize two ways through which ICL leverages demonstrations. Task recognition (TR) captures the extent to which LLMs can recognize a task through demonstrations -- even without ground-truth labels -- and apply their pre-trained priors, whereas task learning (TL) is the ability to capture new input-label mappings unseen in pre-training. Using a wide range of classification datasets and three LLM families (GPT-3, LLaMA and OPT), we design controlled experiments to disentangle the roles of TR and TL in ICL. We show that (1) models can achieve non-trivial performance with only TR, and TR does not further improve with larger models or more demonstrations; (2) LLMs acquire TL as the model scales, and TL's performance consistently improves with more demonstrations in context. Our findings unravel two different forces behind ICL and we advocate for discriminating them in future ICL research due to their distinct nature.

cs.CL