SearcharxivSearch

arXiv · 2606.00415

Vision-Language Model Ensembles Achieve Human-Expert Accuracy for Galaxy Merger Classification

Abstract

We present a proof-of-concept study demonstrating that an ensemble of Vision--Language Models (VLMs) combined using a Bayesian statistical framework can classify galaxy merger morphologies with accuracy comparable to trained human experts. We deploy 15 VLM classifier configurations, spanning four model architectures (Gemma-4 E2B, Gemma-4 E4B, Qwen2.5-VL, and Qwen3-VL) tested with up to four prompt engineering strategies each. We evaluate their performance against a truth-known sample of 41 VELA+SUNRISE mock galaxy images from Lambrides et al. 2021. The VLM ensemble achieves 83.3\% accuracy on confident classifications (merger probability $p_{\rm M} \ge 0.8$ or $p_{\rm M} \le 0.2$), with 5 misclassified galaxies. The ensemble recovers the population merger fraction to within $0.66\sigma$ of the truth ($f_{\rm M} = 0.52 \pm 0.09$ vs.\ true value of 0.585). Bayesian weighting improves overall accuracy by 17.1 percentage points over simple majority voting, with sensitivity improving by 29.2 percentage points. The VLM ensemble produces 5 misclassified galaxies (2 false positives, 3 false negatives), comparable to the 6 misclassifications (5 false positives, 1 false negative) reported for human classifiers by Lambrides et al. 2021. The apparent differences in error profiles are not statistically significant given the sample size of 41 galaxies. VLMs also produce more moderate per-galaxy merger probability distributions (27\% uncertain) than the more polarized human distributions (15\% uncertain), though this difference is also consistent with statistical fluctuation. These results establish VLMs as scalable, reproducible alternatives to human classifiers within a Bayesian probabilistic merger-fraction framework, with direct applications to large galaxy samples from current and future surveys.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marco Chiaberge, Elias Stengel-Eskin, Massimo Stiavelli, Colin Norman. 2026-05-29. Vision-Language Model Ensembles Achieve Human-Expert Accuracy for Galaxy Merger Classification. https://arxiv.org/abs/2606.00415

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Fast Dynamical Modelling of Milky Way Globular Clusters -- II. Impacts of Black Hole Prescriptions

The populations of stellar-mass black holes (BHs) in globular clusters (GCs) play a key role in their dynamical evolution, however the mechanisms surrounding their formation and retention are uncertain. In this work, we extend the analysis of Paper I by fitting coupled rapid cluster evolution and multimass equilibrium models to a large sample of Milky Way GCs, under a variety of prescriptions for stellar evolution, BH formation and supernovae (SN) natal kicks. We explore the impacts of adopting SSE or PARSEC (through SEVN) prescriptions for BH initial-final mass relations, the rapid or delayed SN fallback mechanisms, and an ad hoc grid of kick strengths ejecting between 40 and 80 per cent of all BHs formed. All models reproduce the same present-day conditions despite starting from notably different initial BH populations, due to the correlation found between the initial cluster densities and initial BH mass fractions. A linear relationship is found between the (log) initial half-mass density and the initial BH mass fraction, with the SEVN models resulting in median densities ($\rho_{h,0} \sim 10^{7.2\pm1.1}\,{M_\odot pc^{-3}}$) nearly an order of magnitude higher than those of SSE ($\rho_{h,0} \sim 10^{6.4\pm0.9}\,{M_\odot pc^{-3}}$). We also find that both the bottom-light initial mass functions and the present-day BH mass fractions previously inferred are relatively robust against the stellar evolution models and natal kick prescriptions assumed. Finally, we discuss the implications of these results on the expected numbers and properties of dynamical binary-BH mergers, and the growth of intermediate-mass BHs.

astro-ph.GA

SPURS: An Ultra-deep View Inside the Compact, Nitrogen-Enriched Nuclei of Little Red Dots

We present the first ultra-deep rest-UV spectroscopy of four UV-bright Little Red Dots (LRDs), obtained from the SPURS Cycle 4 Large Program. The spectra reveal broad CIV (FWHM $\approx2700-2800$ km s$^{-1}$) in two LRDs, alongside narrow-line densities elevated above star-forming galaxies ($n_e\sim10^4-10^5$ cm$^{-3}$, reaching $10^6$ cm$^{-3}$ in the most extreme source) and nitrogen-enhancements in all four LRDs. We detect broad HeII emission (FWHM $\approx930$ km s$^{-1}$) in one LRD, and two others with fast P-Cygni absorption ($\gtrsim2200$ km s$^{-1}$). Strong interstellar absorption lines and Ly$\alpha$ damping wings reveal the UV continuum is deeply embedded in neutral gas ($N_{\rm HI}\gtrsim10^{22}$ cm$^{-2}$) in all four LRDs. Detections of fluorescent FeII and OI emission and fine-structure absorption indicate this gas lies close to the UV-emitting region. In archival $z>4$ samples, we find nitrogen and strong CIII] emission are significantly more common in LRDs than in the galaxy population. The transmission of broad CIV, tracing the broad-line region or cocoon, depends on rest-optical color within our sample, consistent with an orientation-dependent picture in which bluer, less obscured sightlines offer a more direct, polar view of the central engine and its outflows. We find several potential signatures of very massive stars, whose winds may contribute to nitrogen enhancement. We investigate other abundance patterns expected from supermassive stars but our results are inconclusive. Our results place the UV-emitting region within $\lesssim8$ pc of the LRD nucleus, consistent with an actively assembling nuclear star cluster. Dynamical interactions in this extremely dense environment, including tidal disruption of stars, may explain the high incidence of nitrogen enhancements in LRDs.

astro-ph.GA

Nitrogen-Loud Quasars from the Dark Energy Spectroscopic Instrument. I. Sample Selection and Basic Properties

We present the largest sample to date of nitrogen-loud (N-loud) quasars with strong broad N IV] $\lambda1486$ and/or N III] $\lambda1750$ emission lines over the redshift range $1.6 < z < 4.3$, selected from the Dark Energy Spectroscopic Instrument (DESI) Data Release 1. The final sample contains 1,993 N-loud quasars, corresponding to about 1.2% of the parent quasar sample. The $L_{1450}$ distribution of the N-loud quasars is broadly similar to that of the DESI parent sample, but their redshift distribution is distinct, with a stronger concentration around $z \sim 2.5$--3. Their composite spectrum displays a broadly similar UV continuum shape to that of the parent quasars, while showing significantly enhanced broad nitrogen emission features, including N V, N IV], and N III]. Other metal emission features also show a moderate enhancement. Relative to a control sample matched in redshift and UV continuum luminosity, the N-loud quasars show systematically narrower broad C IV and Mg II emission lines, lower single-epoch virial black hole masses, and higher Eddington ratios, suggesting that N-loud quasars may preferentially appear during a relatively rapid black hole accretion phase. The radio-loud fraction is 10.1%, with the highest fraction among objects exhibiting both N III] and N IV] emission. The catalog provides a statistical baseline for future studies of nitrogen enhancement and its physical origin.

astro-ph.GA