SearcharxivSearch

arXiv subjects

Kelly Lockhart

Publications and source records attributed to Kelly Lockhart.

7 recordsLinked to original sources

How to Craft the Right Language AI Policy For Your Research Group (Some Assembly Required)

Language AI is rapidly becoming part of the astronomy research ecosystem, prompting research teams to develop policies governing its use. But resources and advice for AI adoption assume that all research groups share the same goals and values. This paper lays out an argument for why there is no single "correct" AI policy for astronomy research groups. Instead, we introduce four research laboratory archetypes with competing research priorities, and we use them to explore how a small laboratory or research group's priorities shape decisions about AI's impact on research productivity, scientist development, scientific integrity, and data governance. Rather than prescribing a universal set of rules, this paper provides a framework for aligning AI policies with a group's scientific values and mission. The objective is to help research leaders decide how Language AI should be used within their particular research environment.

astro-ph.IM

AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics

Scientific multi-label text classification suffers from extreme class imbalance, where specialized terminology exhibits severe power-law distributions that challenge standard classification approaches. Existing scientific corpora lack comprehensive controlled vocabularies, focusing instead on broad categories and limiting systematic study of extreme imbalance. We introduce AstroConcepts, a corpus of English abstracts from 21,702 published astrophysics papers, labeled with 2,367 concepts from the Unified Astronomy Thesaurus. The corpus exhibits severe label imbalance, with 76% of concepts having fewer than 50 training examples. By releasing this resource, we enable systematic study of extreme class imbalance in scientific domains and establish strong baselines across traditional, neural, and vocabulary-constrained LLM methods. Our evaluation reveals three key patterns that provide new insights into scientific text classification. First, vocabulary-constrained LLMs achieve competitive performance relative to domain-adapted models in astrophysics classification, suggesting a potential for parameter-efficient approaches. Second, domain adaptation yields relatively larger improvements for rare, specialized terminology, although absolute performance remains limited across all methods. Third, we propose frequency-stratified evaluation to reveal performance patterns that are hidden by aggregate scores, thereby making robustness assessment central to scientific multi-label evaluation. These results offer actionable insights for scientific NLP and establish benchmarks for research on extreme imbalance.

cs.CL

Do Lexical and Contextual Coreference Resolution Systems Degrade Differently under Mention Noise? An Empirical Study on Scientific Software Mentions

We present our participation in the SOMD 2026 shared task on cross-document software mention coreference resolution, where our systems ranked second across all three subtasks. We compare two fine-tuning-free approaches: Fuzzy Matching (FM), a lexical string-similarity method, and Context Aware Representations (CAR), which combines mention-level and document-level embeddings. Both achieve competitive performance across all subtasks (CoNLL F1 of 0.94-0.96), with CAR consistently outperforming FM by 1 point on the official test set, consistent with the high surface regularity of software names, which reduces the need for complex semantic reasoning. A controlled noise-injection study reveals complementary failure modes: as boundary noise increases, CAR loses only 0.07 F1 points from clean to fully corrupted input, compared to 0.20 for FM, whereas under mention substitution, FM degrades more gracefully (0.52 vs. 0.63). Our inference-time analysis shows that FM scales superlinearly with corpus size, whereas CAR scales approximately linearly, making CAR the more efficient choice at large scale. These findings suggest that system selection should be informed by both the noise profile of the upstream mention detector and the scale of the target corpus. We release our code to support future work on this underexplored task.

cs.CL

INDUS: Effective and Efficient Language Models for Scientific Applications

Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs trained using domain-focused corpora perform better on specialized tasks. Inspired by this insight, we developed INDUS, a comprehensive suite of LLMs tailored for the closely-related domains of Earth science, biology, physics, heliophysics, planetary sciences and astrophysics, and trained using curated scientific corpora drawn from diverse data sources. The suite of models include: (1) an encoder model trained using domain-specific vocabulary and corpora to address NLP tasks, (2) a contrastive-learning based text embedding model trained using a diverse set of datasets to address information retrieval tasks and (3) smaller versions of these models created using knowledge distillation for applications which have latency or resource constraints. We also created three new scientific benchmark datasets, CLIMATE-CHANGE NER (entity-recognition), NASA-QA (extractive QA) and NASA-IR (IR) to accelerate research in these multi-disciplinary fields. We show that our models outperform both general-purpose (RoBERTa) and domain-specific (SCIBERT) encoders on these new tasks as well as existing tasks in the domains of interest. Furthermore, we demonstrate the use of these models in two industrial settings -- as a retrieval model for large-scale vector search applications and in automatic content tagging systems.

cs.CL

A comparison of the morphological properties between local and z~1 infrared luminous galaxies. Are local and high-z (U)LIRGs different?

Ultraluminous and luminous infrared galaxies (ULIRGs and LIRGs) are the most extreme star-forming galaxies in the universe, and dominate the total star formation rate density at z>1. In the local universe (z<0.3), the majority of ULIRGs and a significant portion of LIRGs are triggered by interactions between gas-rich spiral galaxies, yet it is unclear if this is still the case at high-z. To investigate the relative importance of galaxy interactions in infrared luminous galaxies, we carry out a comparison of optical morphological properties between local (U)LIRGs and (U)LIRGs at z=0.5-1.5 based on the same sample selection, morphology classification scheme, and optical morphology at similar rest-frame wavelengths. In addition, we quantify the systematics in comparing local and high-z datasets by constructing a redshifted dataset from local (U)LIRGs, in which its data quality mimics the high-z dataset. Based on the Gini-M20 classification scheme, we find that the fraction of interacting systems decreases by ~8% from local to z<~1, and it is consistent with the reduction between local and redshifted datasets (6(+14-6)%). Based on visual classifications, the merger fraction of local ULIRGs is found to be ~20% lower compared to published results, and the reduction due to redshifiting is 15(+10-8)%. Consequently, the differences of merger fractions between local and z<~1 (U)LIRGs is only ~17%. These results demonstrate that there is no strong evolution in the fraction of (U)LIRGs classified as mergers at least out to z~1. At z>1, the morphology types of ~30% of (U)LIRGs can not be determined due to their faintness in the F814W-band, and thus the merger fraction measured at z>1 suffers from large uncertainties.

astro-ph.GA

The role of galaxy interaction in the SFR-M relation: characterizing morphological properties of Herschel-selected galaxies at 0.2<z<1.5

Galaxy interactions/mergers have been shown to dominate the population of IR luminous galaxies (log(LIR)>11.6Lsun) in the local Universe (z<0.25). Recent studies based on the relation between galaxies' star formation rates and stellar mass (the SFR-M relation or the galaxy main sequence (MS)) have suggested that galaxy interaction/mergers may only become significant when galaxies fall well above the galaxy MS. Since the typical SFR at given M increases with redshift, the existence of galaxy MS implies that massive, IR-luminous galaxies at high-z may not necessarily be driven by galaxy interactions. We examine the role of galaxy interactions in the SFR-M relation by carrying out a morphological analysis of 2084 Herschel-selected galaxies at 0.2 < z < 1.5 in the COSMOS field. Herschel-PACS and -SPIRE observations covering the full 2-deg^2 COSMOS field provide one of the largest far-IR selected samples of high-redshift galaxies with well-determined redshifts to date, with sufficient sensitivity at z ~ 1, to sample objects lying on and above the galaxy MS. Using a detailed visual classification scheme, we show that the fraction of "disk galaxies" decreases and the fraction of "irregular" galaxies increases systematically with increasing LIR out to z ~ 1.5 and z ~ 1.0, respectively. At log(LIR) > 11.5 Lsun, >50% of the objects show evident features of strongly interacting/merger systems, where this percentage is similar to the studies of local IR-luminous galaxies. The fraction of interacting/merger systems also systematically increases with the deviation from the SFR-M relation, supporting the view that galaxies fall above the MS are more dominated by mergers than the MS galaxies. Meanwhile, we find that ~18% of massive IR-luminous MS galaxies are classified as interacting systems, where this population may not evolve through the evolutionary track predicted by a simple gas exhaustion model.

astro-ph.CO

Testing Disk-Locking in NGC 2264

We test analytic predictions from different models of magnetospheric accretion, which invoke disk-locking, using stellar and accretion parameters derived from models of low resolution optical spectra of 36 T Tauri stars (TTSs) in NGC 2264 (age~3 Myrs). Little evidence is found for models that assume purely dipolar field geometries; however, strong support is found in the data for a modified version of the X-wind model (Shu et al. 1994) which allows for non-dipolar field geometries. The trapped flux concept in the X-wind model is key to making the analytic predictions which appear supported in the data. By extension, our analysis provides support for the outflows predicted by the X-wind as these also originate in the trapped flux region. In addition, we find no support in the data for accretion powered stellar winds from young stars. By comparing the analysis presented here of NGC 2264 with a similar analysis of stars in Taurus (age~1-2 Myr), we find evidence that the equilibrium interaction between the magnetic field and accretion disk in TTS systems evolves as the stars grow older, perhaps as the result of evolution of the stellar magnetic field geometry. We compare the accretion rates we derive with accretion rates based on U-band excess, finding good agreement. In addition, we use our accretion parameters to determine the relationship between accretion and H-beta luminosity, again finding good agreement with previously published results; however, we also find that care must be used when applying this relationship due to strong chromospheric emission in young stars which can lead to erroneous results in some cases.

astro-ph.SR