SearcharxivSearch

arXiv subjects

Warren R. Morningstar

Publications and source records attributed to Warren R. Morningstar.

8 recordsLinked to original sources

SASSL: Enhancing Self-Supervised Learning via Neural Style Transfer

Existing data augmentation in self-supervised learning, while diverse, fails to preserve the inherent structure of natural images. This results in distorted augmented samples with compromised semantic information, ultimately impacting downstream performance. To overcome this limitation, we propose SASSL: Style Augmentations for Self Supervised Learning, a novel data augmentation technique based on Neural Style Transfer. SASSL decouples semantic and stylistic attributes in images and applies transformations exclusively to their style while preserving content, generating diverse samples that better retain semantic information. SASSL boosts top-1 image classification accuracy on ImageNet by up to 2 percentage points compared to established self-supervised methods like MoCo, SimCLR, and BYOL, while achieving superior transfer learning performance across various datasets. Because SASSL can be performed asynchronously as part of the data augmentation pipeline, these performance impacts can be obtained with no change in pretraining throughput.

cs.CV

PAC$^m$-Bayes: Narrowing the Empirical Risk Gap in the Misspecified Bayesian Regime

The Bayesian posterior minimizes the "inferential risk" which itself bounds the "predictive risk". This bound is tight when the likelihood and prior are well-specified. However since misspecification induces a gap, the Bayesian posterior predictive distribution may have poor generalization performance. This work develops a multi-sample loss (PAC$^m$) which can close the gap by spanning a trade-off between the two risks. The loss is computationally favorable and offers PAC generalization guarantees. Empirical study demonstrates improvement to the predictive distribution.

cs.LG

Automatic Differentiation Variational Inference with Mixtures

Automatic Differentiation Variational Inference (ADVI) is a useful tool for efficiently learning probabilistic models in machine learning. Generally approximate posteriors learned by ADVI are forced to be unimodal in order to facilitate use of the reparameterization trick. In this paper, we show how stratified sampling may be used to enable mixture distributions as the approximate posterior, and derive a new lower bound on the evidence analogous to the importance weighted autoencoder (IWAE). We show that this "SIWAE" is a tighter bound than both IWAE and the traditional ELBO, both of which are special instances of this bound. We verify empirically that the traditional ELBO objective disfavors the presence of multimodal posterior distributions and may therefore not be able to fully capture structure in the latent space. Our experiments show that using the SIWAE objective allows the encoder to learn more complex distributions which regularly contain multimodality, resulting in higher accuracy and better calibration in the presence of incomplete, limited, or corrupted data.

cs.LG

Density of States Estimation for Out-of-Distribution Detection

Perhaps surprisingly, recent studies have shown probabilistic model likelihoods have poor specificity for out-of-distribution (OOD) detection and often assign higher likelihoods to OOD data than in-distribution data. To ameliorate this issue we propose DoSE, the density of states estimator. Drawing on the statistical physics notion of ``density of states,'' the DoSE decision rule avoids direct comparison of model probabilities, and instead utilizes the ``probability of the model probability,'' or indeed the frequency of any reasonable statistic. The frequency is calculated using nonparametric density estimators (e.g., KDE and one-class SVM) which measure the typicality of various model statistics given the training data and from which we can flag test points with low typicality as anomalous. Unlike many other methods, DoSE requires neither labeled data nor OOD examples. DoSE is modular and can be trivially applied to any existing, trained model. We demonstrate DoSE's state-of-the-art performance against other unsupervised OOD detectors on previously established ``hard'' benchmarks.

cs.LG

Source structure and molecular gas properties from high-resolution CO imaging of SPT-selected dusty star-forming galaxies

We present Atacama Large Millimeter/submillimeter Array (ALMA) observations of high-J CO lines ($J_\mathrm{up}=6$, 7, 8) and associated dust continuum towards five strongly lensed, dusty, star-forming galaxies (DSFGs) at redshift $z = 2.7$-5.7. These galaxies, discovered in the South Pole Telescope survey, are observed at $0.2''$-$0.4''$ resolution with ALMA. Our high-resolution imaging coupled with the lensing magnification provides a measurement of the structure and kinematics of molecular gas in the background galaxies with spatial resolutions down to kiloparsec scales. We derive visibility-based lens models for each galaxy, accurately reproducing observations of four of the galaxies. Of these four targets, three show clear velocity gradients, of which two are likely rotating disks. We find that the reconstructed region of CO emission is less concentrated than the region emitting dust continuum even for the moderate-excitation CO lines, similar to what has been seen in the literature for lower-excitation transitions. We find that the lensing magnification of a given source can vary by 20-50% across the line profile, between the continuum and line, and between different CO transitions. We apply Large Velocity Gradient (LVG) modeling using apparent and intrinsic line ratios between lower-J and high-J CO lines. Ignoring these magnification variations can bias the estimate of physical properties of interstellar medium of the galaxies. The magnitude of the bias varies from galaxy to galaxy and is not necessarily predictable without high resolution observations.

astro-ph.GA

Data-Driven Reconstruction of Gravitationally Lensed Galaxies using Recurrent Inference Machines

We present a machine learning method for the reconstruction of the undistorted images of background sources in strongly lensed systems. This method treats the source as a pixelated image and utilizes the Recurrent Inference Machine (RIM) to iteratively reconstruct the background source given a lens model. Our architecture learns to minimize the likelihood of the model parameters (source pixels) given the data using the physical forward model (ray tracing simulations) while implicitly learning the prior of the source structure from the training data. This results in better performance compared to linear inversion methods, where the prior information is limited to the 2-point covariance of the source pixels approximated with a Gaussian form, and often specified in a relatively arbitrary manner. We combine our source reconstruction network with a convolutional neural network that predicts the parameters of the mass distribution in the lensing galaxies directly from telescope images, allowing a fully automated reconstruction of the background source images and the foreground mass distribution.

astro-ph.IM

Analyzing interferometric observations of strong gravitational lenses with recurrent and convolutional neural networks

We use convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to estimate the parameters of strong gravitational lenses from interferometric observations. We explore multiple strategies and find that the best results are obtained when the effects of the dirty beam are first removed from the images with a deconvolution performed with an RNN-based structure before estimating the parameters. For this purpose, we use the recurrent inference machine (RIM) introduced in Putzky & Welling (2017). This provides a fast and automated alternative to the traditional CLEAN algorithm. We obtain the uncertainties of the estimated parameters using variational inference with Bernoulli distributions. We test the performance of the networks with a simulated test dataset as well as with five ALMA observations of strong lenses. For the observed ALMA data we compare our estimates with values obtained from a maximum-likelihood lens modeling method which operates in the visibility space and find consistent results. We show that we can estimate the lensing parameters with high accuracy using a combination of an RNN structure performing image deconvolution and a CNN performing lensing analysis, with uncertainties less than a factor of two higher than those achieved with maximum-likelihood methods. Including the deconvolution procedure performed by RIM, a single evaluation can be done in about a second on a single GPU, providing a more than six orders of magnitude increase in analysis speed while using about eight orders of magnitude less computational resources compared to maximum-likelihood lens modeling in the uv-plane. We conclude that this is a promising method for the analysis of mm and cm interferometric data from current facilities (e.g., ALMA, JVLA) and future large interferometric observatories (e.g., SKA), where an analysis in the uv-plane could be difficult or unfeasible.

astro-ph.IM

The Spin of the Black Hole GS 1124-683: Observation of a Retrograde Accretion Disk?

We re-examine archival Ginga data for the black hole binary system GS 1124-683, obtained when the system was undergoing its 1991 outburst. Our analysis estimates the dimensionless spin parameter a=cJ/GM^2 by fitting the X-ray continuum spectra obtained while the system was in the "Thermal Dominant" state. For likely values of mass and distance, we find the spin to be a=-0.25 (-0.64, +0.05) (90% confidence), implying that the disk is retrograde (i.e. rotating antiparallel to the spin axis of the black hole). We note that this measurement would be better constrained if the distance to the binary and the mass of the black hole were more accurately determined. This result is unaffected by the model used to fit the hard component of the spectrum. In order to be able to recover a prograde spin, the mass of the black hole would need to be at least 15.25 Msun, or the distance would need to be less than 4.5 kpc, both of which disagree with previous determinations of the black hole mass and distance. If we allow f_col to be free, we obtain no useful spin constraint. We discuss our results in the context of recent spin measurements and implications for jet production.

astro-ph.HE