SearcharxivSearch

arXiv subjects

David W. Hogg

Publications and source records attributed to David W. Hogg.

At least 19 recordsLinked to original sources

Robust Heteroskedastic Matrix Factorization: A Generalization of PCA that Flags Outliers and Handles Missing Data

We present Robust Heteroskedastic Matrix Factorization (RHMF), a generalization of Principal Component Analysis (PCA) that is robust to outliers, handles per-feature uncertainties and missing data, and automatically flags per-feature and per-object anomalies. RHMF is useful both in recovering a low-dimensional embedding unspoiled by bad data or anomalies, and in identifying those anomalies. It utilises an iterative reweighting algorithm that implicitly maximizes a Student-t likelihood. This admits an equivalent probabilistic interpretation as fitting a hierarchical model with per-data-point latent variances. We deliver a fast JAX implementation, Robusta-HMF, and practical guidance for users. We demonstrate the ability of the model to identify and mitigate outliers of different classes. Identification accuracy is contingent on the choice of hyperparameters, but we show that these can be set reliably by cross-validation. We also apply RHMF to RVS spectra from Gaia DR3 to find main-sequence stars that are strange relative to their neighbors in color-magnitude space. We highlight specific examples, including a known binary hosting a Be star, and M-dwarfs with subtle emission in the Ca II triplet lines, indicative of accretion or magnetic activity, which would not be obvious to identify by eye.

astro-ph.IM

Learning What's Real: Disentangling Signal and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics

Data collected from the physical world is always a combination of multiple sources: an underlying signal from the physical process of interest and a signal from measurement-dependent artifacts from the sensor or instrument. This secondary signal acts as a confounding factor, limiting our ability to extract information about the physics underlying the phenomena we observe. Furthermore, it complicates the combination of observations in heterogeneous or multi-instrument settings. We propose a deep learning framework that leverages overlapping observations, a dual-encoder architecture, and a counterfactual generation objective to disentangle these factors of variation. The resulting representations explicitly separate intrinsic signals from sensor-specific distortions and noise, and can be used for counterfactual view generation, parameter inference unconfounded by measurement distortions, and instrument-independent similarity search. We demonstrate the effectiveness of our approach on astrophysical galaxy images from the DESI Legacy Imaging Survey (Legacy) and the Hyper Suprime-Cam (HSC) Survey as a representative multi-instrument setting. This framework provides a general recipe for scientific and multi-modal self-supervised pretraining: construct training pairs from overlapping observations of the same physical system, treat sensor- or modality-specific effects as augmentations, and learn invariant representations through counterfactual generation.

astro-ph.IM

Why do we do astrophysics?

At time of writing, large language models (LLMs) are beginning to obtain the ability to design, execute, write up, and referee scientific projects on the data-science side of astrophysics. What implications does this have for our profession? In this white paper, I list - and argue for - a set of facts or "points of agreement" about what astrophysics is, or should be; these include considerations of novelty, people-centrism, trust, and (the lack of) clinical value. I then list and discuss every possible benefit that astrophysics can be seen as bringing to us, and to science, and to universities, and to the world; these include considerations of love, weaponry, and personal (and personnel) development. I conclude with a discussion of two possible (extreme and bad) policy recommendations related to the use of LLMs in astrophysics, dubbed "let-them-cook" and "ban-and-punish." I argue strongly against both of these; it is not going to be easy to develop or adopt good moderate policies.

astro-ph.IM

Active galactic nuclei do not exhibit strictly sinusoidal brightness variations

Periodic variability in active galactic nuclei (AGN) light curves has been proposed as a signature of close supermassive black hole (SMBH) binaries. Recently, 181 candidate SMBH binaries were identified in Gaia DR3 based on apparently stable sinusoidal variability in their $\sim$1000-day light curves. By supplementing Gaia photometry with longer-baseline light curves from the Zwicky Transient Facility (ZTF) and the Catalina Real Time Transient Survey (CRTS), we test whether the reported periodic signals persist beyond the Gaia DR3 time window. We find that in all 116 cases with available ZTF data, the Gaia-inferred periodic model fails to predict subsequent variability, which appears stochastic rather than periodic. The periodic candidates thus overwhelmingly appear to be false positives; red noise contamination appears to be the primary source of false detections. We conclude that truly periodic and sinusoidal AGN variability is exceedingly rare, with at most a few in $10^6$ AGN exhibiting it on 100 to 1000 day timescales. Models predict that the Gaia AGN light curve sample should contain dozens of true SMBH binaries with periods within the observational baseline, so the lack of strictly periodic light curves in the sample suggests that most short-period binary AGN do not have light curves dominated by simple sinusoidal periodicity.

astro-ph.GA

A constrained linear model for continuum normalization of stellar spectra

Inferring stellar parameters and chemical abundances by forward modeling stellar spectra usually requires a spectral synthesis code, or an emulator constructed from a curated training set. In these situations continuum normalization is often implemented as a pre-processing step that is independent of stellar parameters. This leads to results that are biased, or inconsistent across signal-to-noise ratios. A more justified approach is to forward model spectra with all nuisances simultaneously, but in practice this can be an expensive or non-convex optimization procedure. Here we describe a constrained linear model that can fit stellar absorption, telluric transmission, the joint continuum-instrument response. Stellar absorption and telluric transmission are each modeled by factorizing a grid of rectified theoretical spectra into two non-negative matrices with a chosen number of basis components. This model characterizes all possible spectra in many fewer parameters than comparable data-driven models. The non-negativity constraint ensures basis vectors are strictly additive, which limits rectified flux to less than or equal to unity, such that we can distinguish normalized spectra from the joint instrument-continuum response. The model requires no initial guess, and the linearity ensures that inference is convex, stable, and fast. This model allows us to reliably fit nuisances (e.g., tellurics, continuum), and is readily extensible to radial velocity and rotational broadening, without any prior knowledge about the fundamental stellar properties. We demonstrate our method by fitting ESO/HARPS high-resolution echelle spectra of BAFGKM-type stars. With repeat observations of $α$-Centauri A we present results that are best in class: consistent across time to 0.2% at S/N ~ 100, and to better than 0.5% at S/N ~ 30.

astro-ph.SR

The Milky Way's circular velocity curve measured using element abundance gradients

Spectroscopic surveys now supply precise stellar label measurements such as element abundances for large samples of stars throughout the Milky Way. These element abundances are known to correlate with orbital actions or other dynamical invariants. We present a new data-driven method for empirically measuring the circular velocity curve of the Galaxy that uses element abundance gradients in the plane of radial kinematics. We use stellar surface abundances from the $\textit{APOGEE}$ survey combined with kinematic data from the $\textit{Gaia}$ mission. Our results confirm the ordered structure of the Milky Way disk in terms of average [Fe/H] and [Mg/Fe] abundance ratios, and suggest that $\langle$[Fe/H]$\rangle$ traces the radial position of stars in the disk, while $\langle$[Mg/Fe]$\rangle$ traces the orbital excursions around this radius. Our method uses the radial orbit structure in the Galaxy to enable an empirical measurement of the circular velocity curve, epicyclic and azimuthal frequencies, and kinematic gradients across the Milky Way disk. From these measurements, we infer a value of the circular velocity curve at the Solar radius of $v_{c,\odot} = 235.3^{+2.8}_{-3.7}$ km s$^{-1}$ using the most constraining abundance ratio, [Mg/Fe]. We also measure the radial and azimuthal frequencies for a circular orbit at the solar radius, $κ_{0,R_\odot}=36.9^{+0.8}_{-1.0}$ km s$^{-1}$ kpc$^{-1}$ and $Ω_{0,R_\odot}=28.5_{-0.1}^{+0.4}$ km s$^{-1}$ kpc$^{-1}$, respectively. These values lead to an estimate of the Oort constants of $A = 16.5^{+0.1}_{-0.1}$ km s$^{-1}$ kpc$^{-1}$ and $B=-11.9^{+0.1}_{-0.3}$ km s$^{-1}$ kpc$^{-1}$. We measure the radial acceleration at the Solar radius to be $(\frac{\partial Φ}{\partial R})_{\odot} = a_{R_\odot}=7.0^{+0.2}_{-0.1}$ pc Myr$^{-2}$.

astro-ph.GA

Group Averaging for Physics Applications: Accuracy Improvements at Zero Training Cost

Many machine learning tasks in the natural sciences are precisely equivariant to particular symmetries. Nonetheless, equivariant methods are often not employed, perhaps because training is perceived to be challenging, or the symmetry is expected to be learned, or equivariant implementations are seen as hard to build. Group averaging is an available technique for these situations. It happens at test time; it can make any trained model precisely equivariant at a (often small) cost proportional to the size of the group; it places no requirements on model structure or training. It is known that, under mild conditions, the group-averaged model will have a provably better prediction accuracy than the original model. Here we show that an inexpensive group averaging can improve accuracy in practice. We take well-established benchmark machine learning models of differential equations in which certain symmetries ought to be obeyed. At evaluation time, we average the models over a small group of symmetries. Our experiments show that this procedure always decreases the average evaluation loss, with improvements of up to 37\% in terms of the VRMSE. The averaging produces visually better predictions for continuous dynamics. This short paper shows that, under certain common circumstances, there are no disadvantages to imposing exact symmetries; the ML4PS community should consider group averaging as a cheap and simple way to improve model accuracy.

cs.LG

Evolved stars with inconsistent age estimates: Abundance outliers or mass transfer products?

In the Milky Way disk there is a strong trend linking stellar age to surface element abundances. Here we explore this relationship with a dataset of 8,803 red-giant and red-clump stars with both asteroseismic data from NASA Kepler Mission and surface abundances from the SDSS-V MWM. We find, with a k-nearest-neighbors approach, that the [Mg/H] and [Fe/Mg] abundance ratios predict asteroseismic ages to an accuracy of about 2 Gyr for the majority of stars. That said, there are substantial outlier stars whose surface abundances do not match their asteroseismic ages. Because asteroseismic ages for these stars are fundamentally based on density or mass, these outliers are mass-transfer candidates. Stars whose surface abundances predict a younger age (higher mass) than what's seen in the asteroseismology are mass accretor candidates (MAC); stars whose abundances predict an older age (lower mass) than the asteroseismology age are mass donor candidates (MDC). We create precise control samples, matched according to (1) surface abundances and (2) asteroseismic ages, for both the MAC and MDC stars; we use these to find slight differences in rotational velocity, [C/N], and [Na/Mg] between the mass-transfer candidates and their abundance neighbors. We find no drastic differences in kinematics, orbital invariants, UV excess, or other stellar abundances between outliers and their abundance neighbors. We deliver 377 mass-transfer candidates for follow-up observations. This project implicitly suggests a fundamental limit on the reliability of asteroseismic ages, and supports existing evidence that age-abundance outliers are products of binary mass transfer.

astro-ph.SR

Optical Spectroscopy Reveals Hidden Neutron-capture Elemental Abundance Differences among APOGEE-identified Chemical Doppelgängers

Grouping stars by chemical similarity has the potential to reveal the Milky Way's evolutionary history. The APOGEE stellar spectroscopic survey has the resolution and sensitivity for this task. However, APOGEE lacks access to strong lines of neutron-capture elements ($Z > 28$) which have nucleosynthetic origins that are distinct from those of the lighter elements. We assess whether APOGEE abundances are sufficient for selecting chemically similar disk stars by identifying 25 pairs of chemical ``doppelgangers'' in APOGEE DR17 and following them up with the Tull spectrograph, an optical, $R \sim 60{,}000$ echelle on the McDonald Observatory 2.7-m telescope. Line-by-line differential analyses of pairs' optical spectra reveals neutron-capture (Y, Zr, Ba, La, Ce, Nd, and Eu) elemental abundance differences of $Δ$[X/Fe] $\rm \sim 0.020 \pm 0.015$ to $0.380 \pm 0.15$ dex (4--140%), and up to 0.05 dex (12%) on average, a factor of 1--2 times higher than intra-cluster pairs. This is despite the pairs sharing nearly identical APOGEE-reported abundances and [C/N] ratios, a tracer of giant-star age. This work illustrates that even when APOGEE abundances derived from SNR $> 300$ spectra are available, optically-measured neutron-capture element abundances contain critical information about composition similarity. These results hold implications for the chemical dimensionality of the disk, mixing within the interstellar medium, and chemical tagging with the neutron-capture elements.

astro-ph.SR

Causal Foundation Models: Disentangling Physics from Instrument Properties

Foundation models for structured time series data must contend with a fundamental challenge: observations often conflate the true underlying physical phenomena with systematic distortions introduced by measurement instruments. This entanglement limits model generalization, especially in heterogeneous or multi-instrument settings. We present a causally-motivated foundation model that explicitly disentangles physical and instrumental factors using a dual-encoder architecture trained with structured contrastive learning. Leveraging naturally occurring observational triplets (i.e., where the same target is measured under varying conditions, and distinct targets are measured under shared conditions) our model learns separate latent representations for the underlying physical signal and instrument effects. Evaluated on simulated astronomical time series designed to resemble the complexity of variable stars observed by missions like NASA's Transiting Exoplanet Survey Satellite (TESS), our method significantly outperforms traditional single-latent space foundation models on downstream prediction tasks, particularly in low-data regimes. These results demonstrate that our model supports key capabilities of foundation models, including few-shot generalization and efficient adaptation, and highlight the importance of encoding causal structure into representation learning for structured data.

cs.LG

$Lux$: A generative, multi-output, latent-variable model for astronomical data with noisy labels

The large volume of spectroscopic data available now and from near-future surveys will enable high-dimensional measurements of stellar parameters and properties. Current methods for determining stellar labels from spectra use physics-driven models, which are computationally expensive and have limitations in their accuracy due to simplifications. While machine learning methods provide efficient paths toward emulating physics-based pipelines, they often do not properly account for uncertainties and have complex model structure, both of which can lead to biases and inaccurate label inference. Here we present $Lux$: a data-driven framework for modeling stellar spectra and labels that addresses prior limitations. $Lux$ is a generative, multi-output, latent variable model framework built on JAX for computational efficiency and flexibility. As a generative model, $Lux$ properly accounts for uncertainties and missing data in the input stellar labels and spectral data and can either be used in probabilistic or discriminative settings. Here, we present several examples of how $Lux$ can successfully emulate methods for precise stellar label determinations for stars ranging in stellar type and signal-to-noise from the $APOGEE$ surveys. We also show how a simple $Lux$ model is successful at performing label transfer between the $APOGEE$ and $GALAH$ surveys. $Lux$ is a powerful new framework for the analysis of large-scale spectroscopic survey data. Its ability to handle uncertainties while maintaining high precision makes it particularly valuable for stellar survey label inference and cross-survey analysis, and the flexible model structure allows for easy extension to other data types.

astro-ph.IM

A formula for the area of a triangle: Useless, but explicitly in Deep Sets form

Any permutation-invariant function of data points $\vec{r}_i$ can be written in the form $ρ(\sum_iϕ(\vec{r}_i))$ for suitable functions $ρ$ and $ϕ$. This form - known in the machine-learning literature as Deep Sets - also generates a map-reduce algorithm. The area of a triangle is a permutation-invariant function of the locations $\vec{r}_i$ of the three corners $1\leq i\leq 3$. We find the polynomial formula for the area of a triangle that is explicitly in Deep Sets form. This project was motivated by questions about the fundamental computational complexity of $n$-point statistics in cosmology; that said, no insights of any kind were gained from these results.

astro-ph.CO

On the effects of parameters on galaxy properties in CAMELS and the predictability of $Ω_{\rm m}$

Recent analyses of cosmological hydrodynamic simulations from CAMELS have shown that machine learning models can predict the parameter describing the total matter content of the universe, $Ω_{\rm m}$, from the features of a single galaxy. We investigate the statistical properties of two of these simulation suites, IllustrisTNG and ASTRID, confirming that $Ω_{\rm m}$ induces a strong displacement on the distribution of galaxy features. We also observe that most other parameters have little to no effect on the distribution, except for the stellar-feedback parameter $A_{SN1}$, which introduces some near-degeneracies that can be broken with specific features. These two properties explain the predictability of $Ω_{\rm m}$. We use Optimal Transport to further measure the effect of parameters on the distribution of galaxy properties, which is found to be consistent with physical expectations. However, we observe discrepancies between the two simulation suites, both in the effect of $Ω_{\rm m}$ on the galaxy properties and in the distributions themselves at identical parameter values. Thus, although $Ω_{\rm m}$'s signature can be easily detected within a given simulation suite using just a single galaxy, applying this result to real observational data may prove significantly more challenging.

astro-ph.CO

Identification of 30,000 White Dwarf-Main Sequence binaries candidates from Gaia DR3 BP/RP(XP) low-resolution spectra

White dwarf-main sequence (WDMS) binary systems are essential probes for understanding binary stellar evolution and play a pivotal role in constraining theoretical models of various transient phenomena. In this study, we construct a catalog of WDMS binaries using Gaia DR3's low-resolution BP/RP (XP) spectra. Our approach integrates a model-independent neural network for spectral modelling with Gaussian Process Classification to accurately identify WDMS binaries among over 10 million stars within 1 kpc. This study identify approximately 30,000 WDMS binary candidates, including ~1,700 high-confidence systems confirmed through spectral fitting. Our technique is shown to be effective at detecting systems where the main-sequence star dominates the spectrum - cases that have historically challenged conventional methods. Validation using GALEX photometry reinforces the reliability of our classifications: 70\% of candidates with an absolute magnitude $M_{G} > 7$ exhibit UV excess, a characteristic signature of white dwarf companions. Our all-sky catalog of WDMS binaries expands the available dataset for studying binary evolution and white dwarf physics and sheds light on the formation of WDMS.

astro-ph.SR

zoomies: A tool to infer stellar age from vertical action in Gaia data

Stellar age measurements are fundamental to understanding a wide range of astronomical processes, including Galactic dynamics, stellar evolution, and planetary system formation. However, extracting age information from main-sequence stars is complicated, with techniques often relying on age proxies in the absence of direct measurements. The Gaia data releases have enabled detailed studies of the dynamical properties of stars within the Milky Way, offering new opportunities to understand the relationship between stellar age and dynamics. In this study, we leverage high-precision astrometric data from Gaia DR3 to construct a stellar age prediction model based only on stellar dynamical properties, namely the vertical action. We calibrate two distinct, hierarchical stellar age--vertical action relations, first employing asteroseismic ages for red-giant-branch stars, then isochrone ages for main-sequence turn-off stars. We describe a framework called "zoomies" based on this calibration, by which we can infer ages for any star given its vertical action. This tool is open-source and intended for community use. We compare dynamical age estimates from "zoomies" with age measurements from open clusters and asteroseismology. We use "zoomies" to generate and compare dynamical age estimates for stars from the Kepler, K2, and TESS exoplanet transit surveys. While dynamical age relations are associated with large uncertainty, they are generally mass independent and depend on homogeneously measured astrometric data. These age predictions are uniquely useful for large-scale demographic investigations, especially in disentangling the relationship between planet occurrence, metallicity, and age for low-mass stars.

astro-ph.SR

A Compact, Coherent Representation of Stellar Surface Variation in the Spectral Domain

Time-varying inhomogeneities on stellar surfaces constitute one of the largest sources of radial velocity (RV) error for planet detection and characterization. We show that stellar variations, because they manifest on coherent, rotating surfaces, give rise to changes that are complex but useably compact and coherent in the spectral domain. Methods for disentangling stellar signals in RV measurements benefit from modeling the full domain of spectral pixels. We simulate spectra of spotted stars using starry and construct a simple spectrum projection space that is sensitive to the orientation and size of stellar surface features. Regressing measured RVs in this projection space reduces RV scatter by 60-80% while preserving planet shifts. We note that stellar surface variability signals do not manifest in spectral changes that are purely orthogonal to a Doppler shift or exclusively asymmetric in line profiles; enforcing orthogonality or focusing exclusively on asymmetric features will not make use of all the information present in the spectra. We conclude with a discussion of existing and possible implementations on real data based on the presented compact, coherent framework for stellar signal mitigation.

astro-ph.SR

Equivariant geometric convolutions for emulation of dynamical systems

Machine learning methods are increasingly being employed as surrogate models in place of computationally expensive and slow numerical integrators for a bevy of applications in the natural sciences. However, while the laws of physics are relationships between scalars, vectors, and tensors that hold regardless of the frame of reference or chosen coordinate system, surrogate machine learning models are not coordinate-free by default. We enforce coordinate freedom by using geometric convolutions in three model architectures: a ResNet, a Dilated ResNet, and a UNet. In numerical experiments emulating 2D compressible Navier-Stokes, we see better accuracy and improved stability compared to baseline surrogate models in almost all cases. The ease of enforcing coordinate freedom without making major changes to the model architecture provides an exciting recipe for any CNN-based method applied to an appropriate class of problems

cs.LG

Many elements matter: Detailed abundance patterns reveal star-formation and enrichment differences among Milky Way structural components

Many nucleosynthetic channels create the elements, but two-parameter models characterized by $α$ and Fe nonetheless predict stellar abundances in the Galactic disk to accuracies of 0.02 to 0.05 dex for most measured elements, near the level of current abundance uncertainties. It is difficult to make individual measurements more precise than this to investigate lower-amplitude nucleosynthetic effects, but population studies of mean abundance patterns can reveal more subtle abundance differences. Here we look at the detailed abundances for 67315 stars from APOGEE DR17, but in the form of abundance residuals away from a best-fit two-parameter, data-driven nucleosynthetic model. We find that these residuals show complex structures with respect to age, guiding radius, and vertical action that are not random and are also not strongly correlated with sources of systematic error such as surface gravity, effective temperature, and radial velocity. The residual patterns, especially in Na, C+N, Ni, Mn, and Ce, trace kinematic structures in the Milky Way, such as the inner disk, thick disk, and flared outer disk. A principal component analysis suggests that most of the observed structure is low-dimensional and can be explained by a few eigenvectors. We find that some, but not all, of the effects in the low-$α$ disk can be explained by dilution with fresh gas, so that abundance ratios resemble those of stars with higher metallicity. The patterns and maps we provide could be combined with accurate forward models of nucleosynthesis, star formation, and gas infall to provide a more detailed picture of star and element formation in different Milky Way components.

astro-ph.GA