SearcharxivSearch

arXiv subjects

Richard Scalzo

Publications and source records attributed to Richard Scalzo.

At least 19 recordsLinked to original sources

Emergence of Computational Structure in a Neural Network Physics Simulator

Neural networks often have identifiable computational structures - components of the network which perform an interpretable algorithm or task - but the mechanisms by which these emerge and the best methods for detecting these structures are not well understood. In this paper we investigate the emergence of computational structure in a transformer-like model trained to simulate the physics of a particle system, where the transformer's attention mechanism is used to transfer information between particles. We show that (a) structures emerge in the attention heads of the transformer which learn to detect particle collisions, (b) the emergence of these structures is associated to degenerate geometry in the loss landscape, and (c) the dynamics of this emergence follows a power law. This suggests that these components are governed by a degenerate "effective potential". These results have implications for the convergence time of computational structure within neural networks and suggest that the emergence of computational structure can be detected by studying the dynamics of network components.

cs.LG

Semi-Supervised Learning for Lensed Quasar Detection

Lensed quasars are key to many areas of study in astronomy, offering a unique probe into the intermediate and far universe. However, finding lensed quasars has proved difficult despite significant efforts from large collaborations. These challenges have limited catalogues of confirmed lensed quasars to the hundreds, despite theoretical predictions that they should be many times more numerous. We train machine learning classifiers to discover lensed quasar candidates. By using semi-supervised learning techniques we leverage the large number of potential candidates as unlabelled training data alongside the small number of known objects, greatly improving model performance. We present our two most successful models: (1) a variational autoencoder trained on millions of quasars to reduce the dimensionality of images for input to a dense neural network classifier that can make accurate predictions and (2) a convolutional neural network trained on a mix of labelled and unlabelled data via virtual adversarial training. These models are both capable of producing high-quality candidates, as evidenced by our discovery of GRALJ140833.73+042229.98. The success of our classifier, which uses only multi-band images, is particularly exciting as it can be combined with existing classifiers, which use other data than images, to improve the classifications of both models and discover more lensed quasars.

astro-ph.IM

PINN-Ray: A Physics-Informed Neural Network to Model Soft Robotic Fin Ray Fingers

Modelling complex deformation for soft robotics provides a guideline to understand their behaviour, leading to safe interaction with the environment. However, building a surrogate model with high accuracy and fast inference speed can be challenging for soft robotics due to the nonlinearity from complex geometry, large deformation, material nonlinearity etc. The reality gap from surrogate models also prevents their further deployment in the soft robotics domain. In this study, we proposed a physics-informed Neural Networks (PINNs) named PINN-Ray to model complex deformation for a Fin Ray soft robotic gripper, which embeds the minimum potential energy principle from elastic mechanics and additional high-fidelity experimental data into the loss function of neural network for training. This method is significant in terms of its generalisation to complex geometry and robust to data scarcity as compared to other data-driven neural networks. Furthermore, it has been extensively evaluated to model the deformation of the Fin Ray finger under external actuation. PINN-Ray demonstrates improved accuracy as compared with Finite element modelling (FEM) after applying the data assimilation scheme to treat the sim-to-real gap. Additionally, we introduced our automated framework to design, fabricate soft robotic fingers, and characterise their deformation by visual tracking, which provides a guideline for the fast prototype of soft robotics.

cs.RO

Observing the Galactic Underworld: Predicting photometry and astrometry from compact remnant microlensing events

Isolated black holes (BHs) and neutron stars (NSs) are largely undetectable across the electromagnetic spectrum. For this reason, our only real prospect of observing these isolated compact remnants is via microlensing; a feat recently performed for the first time. However, characterisation of the microlensing events caused by BHs and NSs is still in its infancy. In this work, we perform N-body simulations to explore the frequency and physical characteristics of microlensing events across the entire sky. Our simulations find that every year we can expect $88_{-6}^{+6}$ BH, $6.8_{-1.6}^{+1.7}$ NS and $20^{+30}_{-20}$ stellar microlensing events which cause an astrometric shift larger than 2~mas. Similarly, we can expect $21_{-3}^{+3}$ BH, $18_{-3}^{+3}$ NS and $7500_{-500}^{+500}$ stellar microlensing events which cause a bump magnitude larger than 1~mag. Leveraging a more comprehensive dynamical model than prior work, we predict the fraction of microlensing events caused by BHs as a function of Einstein time to be smaller than previously thought. Comparison of our microlensing simulations to events in Gaia finds good agreement. Finally, we predict that in the combination of Gaia and GaiaNIR data there will be $14700_{-900}^{+600}$ BH and $1600_{-200}^{+300}$ NS events creating a centroid shift larger than 1~mas and $330_{-120}^{+100}$ BH and $310_{-100}^{+110}$ NS events causing bump magnitudes $> 1$. Of these, $<10$ BH and $5_{-5}^{+10}$ NS events should be detectable using current analysis techniques. These results inform future astrometric mission design, such as GaiaNIR, as they indicate that, compared to stellar events, there are fewer observable BH events than previously thought.

astro-ph.GA

Nonlinear wavefront reconstruction from a pyramid sensor using neural networks

The pyramid wavefront sensor (PyWFS) has become increasingly popular to use in adaptive optics (AO) systems due to its high sensitivity. The main drawback of the PyWFS is that it is inherently nonlinear, which means that classic linear wavefront reconstruction techniques face a significant reduction in performance at high wavefront errors, particularly when the pyramid is unmodulated. In this paper, we consider the potential use of neural networks (NNs) to replace the widely used matrix vector multiplication (MVM) control. We aim to test the hypothesis that the neural network (NN)'s ability to model nonlinearities will give it a distinct advantage over MVM control. We compare the performance of a MVM linear reconstructor against a dense NN, using daytime data acquired on the Subaru Coronagraphic Extreme Adaptive Optics system (SCExAO) instrument. In a first set of experiments, we produce wavefronts generated from 14 Zernike modes and the PyWFS responses at different modulation radii (25, 50, 75, and 100 mas). We find that the NN allows for a far more precise wavefront reconstruction at all modulations, with differences in performance increasing in the regime where the PyWFS nonlinearity becomes significant. In a second set of experiments, we generate a dataset of atmosphere-like wavefronts, and confirm that the NN outperforms the linear reconstructor. The SCExAO real-time computer software is used as baseline for the latter. These results suggest that NNs are well positioned to improve upon linear reconstructors and stand to bring about a leap forward in AO performance in the near future.

astro-ph.IM

Learning the Lantern: Neural network applications to broadband photonic lantern modelling

Photonic lanterns allow the decomposition of highly multimodal light into a simplified modal basis such as single-moded and/or few-moded. They are increasingly finding uses in astronomy, optics and telecommunications. Calculating propagation through a photonic lantern using traditional algorithms takes $\sim 1$ hour per simulation on a modern CPU. This paper demonstrates that neural networks can bridge the disparate opto-electronic systems, and when trained can achieve a speed-up of over 5 orders of magnitude. We show that this approach can be used to model photonic lanterns with manufacturing defects as well as successfully generalising to polychromatic data. We demonstrate two uses of these neural network models, propagating seeing through the photonic lantern as well as performing global optimisation for purposes such as photonic lantern funnels and photonic lantern nullers.

physics.optics

SkyMapper Optical Follow-up of Gravitational Wave Triggers: Alert Science Data Pipeline and LIGO/Virgo O3 Run

We present an overview of the SkyMapper optical follow-up program for gravitational-wave event triggers from the LIGO/Virgo observatories, which aims at identifying early GW170817-like kilonovae out to $\sim 200$ Mpc distance. We describe our robotic facility for rapid transient follow-up, which can target most of the sky at $δ<+10°$ to a depth of $i_\mathrm{AB}\approx 20$ mag. We have implemented a new software pipeline to receive LIGO/Virgo alerts, schedule observations and examine the incoming real-time data stream for transient candidates. We adopt a real-bogus classifier using ensemble-based machine learning techniques, attaining high completeness ($\sim$98%) and purity ($\sim$91%) over our whole magnitude range. Applying further filtering to remove common image artefacts and known sources of transients, such as asteroids and variable stars, reduces the number of candidates by a factor of more than 10. We demonstrate the system performance with data obtained for GW190425, a binary neutron star merger detected during the LIGO/Virgo O3 observing campaign. In time for the LIGO/Virgo O4 run, we will have deeper reference images allowing transient detection to $i_\mathrm{AB}\approx $21 mag.

astro-ph.IM

Periodic Astrometric Signal Recovery through Convolutional Autoencoders

Astrometric detection involves a precise measurement of stellar positions, and is widely regarded as the leading concept presently ready to find earth-mass planets in temperate orbits around nearby sun-like stars. The TOLIMAN space telescope[39] is a low-cost, agile mission concept dedicated to narrow-angle astrometric monitoring of bright binary stars. In particular the mission will be optimised to search for habitable-zone planets around Alpha Centauri AB. If the separation between these two stars can be monitored with sufficient precision, tiny perturbations due to the gravitational tug from an unseen planet can be witnessed and, given the configuration of the optical system, the scale of the shifts in the image plane are about one millionth of a pixel. Image registration at this level of precision has never been demonstrated (to our knowledge) in any setting within science. In this paper we demonstrate that a Deep Convolutional Auto-Encoder is able to retrieve such a signal from simplified simulations of the TOLIMAN data and we present the full experimental pipeline to recreate out experiments from the simulations to the signal analysis. In future works, all the more realistic sources of noise and systematic effects present in the real-world system will be injected into the simulations.

astro-ph.IM

Bayesreef: A Bayesian inference framework for modelling reef growth in response to environmental change and biological dynamics

Estimating the impact of environmental processes on vertical reef development in geological time is a very challenging task. pyReef-Core is a deterministic carbonate stratigraphic forward model designed to simulate the key biological and environmental processes that determine vertical reef accretion and assemblage changes in fossil reef drill cores. We present a Bayesian framework called Bayesreef for the estimation and uncertainty quantification of parameters in pyReef-Core that represent environmental conditions affecting the growth of coral assemblages on geological timescales. We demonstrate the existence of multimodal posterior distributions and investigate the challenges of sampling using Markov chain Monte-Carlo (MCMC) methods, which includes parallel tempering MCMC. We use synthetic reef-core to investigate fundamental issues and then apply the methodology to a selected reef-core from the Great Barrier Reef in Australia. The results show that Bayesreef accurately estimates and provides uncertainty quantification of the selected parameters that represent the environment and ecological conditions in pyReef-Core. Bayesreef provides insights into the complex posterior distributions of parameters in pyReef-Core, which provides the groundwork for future research in this area.

stat.AP

The SAMI Galaxy Survey: Bayesian Inference for Gas Disk Kinematics using a Hierarchical Gaussian Mixture Model

We present a novel Bayesian method, referred to as Blobby3D, to infer gas kinematics that mitigates the effects of beam smearing for observations using Integral Field Spectroscopy (IFS). The method is robust for regularly rotating galaxies despite substructure in the gas distribution. Modelling the gas substructure within the disk is achieved by using a hierarchical Gaussian mixture model. To account for beam smearing effects, we construct a modelled cube that is then convolved per wavelength slice by the seeing, before calculating the likelihood function. We show that our method can model complex gas substructure including clumps and spiral arms. We also show that kinematic asymmetries can be observed after beam smearing for regularly rotating galaxies with asymmetries only introduced in the spatial distribution of the gas. We present findings for our method applied to a sample of 20 star-forming galaxies from the SAMI Galaxy Survey. We estimate the global H$α$ gas velocity dispersion for our sample to be in the range $\barσ_v \sim $[7, 30] km s$^{-1}$. The relative difference between our approach and estimates using the single Gaussian component fits per spaxel is $Δ\barσ_v / \barσ_v = - 0.29 \pm 0.18$ for the H$α$ flux-weighted mean velocity dispersion.

astro-ph.GA

Efficiency and robustness in Monte Carlo sampling of 3-D geophysical inversions with Obsidian v0.1.2: Setting up for success

The rigorous quantification of uncertainty in geophysical inversions is a challenging problem. Inversions are often ill-posed and the likelihood surface may be multi-modal; properties of any single mode become inadequate uncertainty measures, and sampling methods become inefficient for irregular posteriors or high-dimensional parameter spaces. We explore the influences of different choices made by the practitioner on the efficiency and accuracy of Bayesian geophysical inversion methods that rely on Markov chain Monte Carlo sampling to assess uncertainty, using a multi-sensor inversion of the three-dimensional structure and composition of a region in the Cooper Basin of South Australia as a case study. The inversion is performed using an updated version of the Obsidian distributed inversion software. We find that the posterior for this inversion has complex local covariance structure, hindering the efficiency of adaptive sampling methods that adjust the proposal based on the chain history. Within the context of a parallel-tempered Markov chain Monte Carlo scheme for exploring high-dimensional multi-modal posteriors, a preconditioned Crank-Nicholson proposal outperforms more conventional forms of random walk. Aspects of the problem setup, such as priors on petrophysics or on 3-D geological structure, affect the shape and separation of posterior modes, influencing sampling performance as well as the inversion results. Use of uninformative priors on sensor noise can improve inversion results by enabling optimal weighting among multiple sensors even if noise levels are uncertain. Efficiency could be further increased by using posterior gradient information within proposals, which Obsidian does not currently support, but which could be emulated using posterior surrogates.

stat.AP

Computer vision-based framework for extracting geological lineaments from optical remote sensing data

The extraction of geological lineaments from digital satellite data is a fundamental application in remote sensing. The location of geological lineaments such as faults and dykes are of interest for a range of applications, particularly because of their association with hydrothermal mineralization. Although a wide range of applications have utilized computer vision techniques, a standard workflow for application of these techniques to mineral exploration is lacking. We present a framework for extracting geological lineaments using computer vision techniques which is a combination of edge detection and line extraction algorithms for extracting geological lineaments using optical remote sensing data. It features ancillary computer vision techniques for reducing data dimensionality, removing noise and enhancing the expression of lineaments. We test the proposed framework on Landsat 8 data of a mineral-rich portion of the Gascoyne Province in Western Australia using different dimension reduction techniques and convolutional filters. To validate the results, the extracted lineaments are compared to our manual photointerpretation and geologically mapped structures by the Geological Survey of Western Australia (GSWA). The results show that the best correlation between our extracted geological lineaments and the GSWA geological lineament map is achieved by applying a minimum noise fraction transformation and a Laplacian filter. Application of a directional filter instead shows a stronger correlation with the output of our manual photointerpretation and known sites of hydrothermal mineralization. Hence, our framework using either filter can be used for mineral prospectivity mapping in other regions where faults are exposed and observable in optical remote sensing data.

cs.CV

SN 2012fr: Ultraviolet, Optical, and Near-Infrared Light Curves of a Type Ia Supernova Observed Within a Day of Explosion

We present detailed ultraviolet, optical and near-infrared light curves of the Type Ia supernova (SN) 2012fr, which exploded in the Fornax cluster member NGC 1365. These precise high-cadence light curves provide a dense coverage of the flux evolution from $-$12 to $+$140 days with respect to the epoch of $B$-band maximum (\tmax). Supplementary imaging at the earliest epochs reveals an initial slow, nearly linear rise in luminosity with a duration of $\sim$2.5 days, followed by a faster rising phase that is well reproduced by an explosion model with a moderate amount of $^{56}$Ni mixing in the ejecta. From an analysis of the light curves, we conclude: $(i)$ explosion occurred $< 22$ hours before the first detection of the supernova, $(ii)$ the rise time to peak bolometric ($λ> 1800 $Å) luminosity was $16.5 \pm 0.6$ days, $(iii)$ the supernova suffered little or no host-galaxy dust reddening, $(iv)$ the peak luminosity in both the optical and near-infrared was consistent with the bright end of normal Type Ia diversity, and $(v)$ $0.60 \pm 0.15 M_{\odot}$ of $^{56}$Ni was synthesized in the explosion. Despite its normal luminosity, SN 2012fr displayed unusually prevalent high-velocity \ion{Ca}{2} and \ion{Si}{2} absorption features, and a nearly constant photospheric velocity of the \ion{Si}{2} $λ$6355 line at $\sim$12,000 \kms\ beginning $\sim$5 days before \tmax. Other peculiarities in the early phase photometry and the spectral evolution are highlighted. SN 2012fr also adds to a growing number of Type Ia supernovae hosted by galaxies with direct Cepheid distance measurements.

astro-ph.HE

The SkyMapper Transient Survey

The SkyMapper 1.3 m telescope at Siding Spring Observatory has now begun regular operations. Alongside the Southern Sky Survey, a comprehensive digital survey of the entire southern sky, SkyMapper will carry out a search for supernovae and other transients. The search strategy, covering a total footprint area of ~2000 deg2 with a cadence of $\leq 5$ days, is optimised for discovery and follow-up of low-redshift type Ia supernovae to constrain cosmic expansion and peculiar velocities. We describe the search operations and infrastructure, including a parallelised software pipeline to discover variable objects in difference imaging; simulations of the performance of the survey over its lifetime; public access to discovered transients; and some first results from the Science Verification data.

astro-ph.IM

The ANU WiFeS SuperNovA Program (AWSNAP)

This paper presents the first major data release and survey description for the ANU WiFeS SuperNovA Program (AWSNAP). AWSNAP is an ongoing supernova spectroscopy campaign utilising the Wide Field Spectrograph (WiFeS) on the Australian National University (ANU) 2.3m telescope. The first and primary data release of this program (AWSNAP-DR1) releases 357 spectra of 175 unique objects collected over 82 equivalent full nights of observing from July 2012 to August 2015. These spectra have been made publicly available via the WISeREP supernova spectroscopy repository. We analyse the AWSNAP sample of Type Ia supernova spectra, including measurements of narrow sodium absorption features afforded by the high spectral resolution of the WiFeS instrument. In some cases we were able to use the integral-field nature of the WiFeS instrument to measure the rotation velocity of the SN host galaxy near the SN location in order to obtain precision sodium absorption velocities. We also present an extensive time series of SN 2012dn, including a near-nebular spectrum which both confirms its "super-Chandrasekhar" status and enables measurement of the sub-solar host metallicity at the SN site.

astro-ph.HE

Measuring nickel masses in Type Ia supernovae using cobalt emission in nebular phase spectra

The light curves of Type Ia supernovae (SNe Ia) are powered by the radioactive decay of $^{56}$Ni to $^{56}$Co at early times, and the decay of $^{56}$Co to $^{56}$Fe from ~60 days after explosion. We examine the evolution of the [Co III] 5892 A emission complex during the nebular phase for SNe Ia with multiple nebular spectra and show that the line flux follows the square of the mass of $^{56}$Co as a function of time. This result indicates both efficient local energy deposition from positrons produced in $^{56}$Co decay, and long-term stability of the ionization state of the nebula. We compile 77 nebular spectra of 25 SN Ia from the literature and present 17 new nebular spectra of 7 SNe Ia, including SN2014J. From these we measure the flux in the [Co III] 5892 A line and remove its well-behaved time dependence to infer the initial mass of $^{56}$Ni ($M_{Ni}$) produced in the explosion. We then examine $^{56}$Ni yields for different SN Ia ejected masses ($M_{ej}$ - calculated using the relation between light curve width and ejected mass) and find the $^{56}$Ni masses of SNe Ia fall into two regimes: for narrow light curves (low stretch s~0.7-0.9), $M_{Ni}$ is clustered near $M_{Ni}$ ~ 0.4$M_\odot$ and shows a shallow increase as $M_{ej}$ increases from ~1-1.4$M_\odot$; at high stretch, $M_{ej}$ clusters at the Chandrasekhar mass (1.4$M_\odot$) while $M_{Ni}$ spans a broad range from 0.6-1.2$M_\odot$. This could constitute evidence for two distinct SN Ia explosion mechanisms.

astro-ph.HE

Ultraviolet Observations of Super-Chandrasekhar Mass Type Ia Supernova Candidates with Swift UVOT

Among Type Ia supernovae (SNe~Ia) exist a class of overluminous objects whose ejecta mass is inferred to be larger than the canonical Chandrasekhar mass. We present and discuss the UV/optical photometric light curves, colors, absolute magnitudes, and spectra of three candidate Super-Chandrasekhar mass SNe--2009dc, 2011aa, and 2012dn--observed with the Swift Ultraviolet/Optical Telescope. The light curves are at the broad end for SNe Ia, with the light curves of SN~2011aa being amongst the broadest ever observed. We find all three to have very blue colors which may provide a means of excluding these overluminous SNe from cosmological analysis, though there is some overlap with the bluest of "normal" SNe Ia. All three are overluminous in their UV absolute magnitudes compared to normal and broad SNe Ia, but SNe 2011aa and 2012dn are not optically overluminous compared to normal SNe Ia. The integrated luminosity curves of SNe 2011aa and 2012dn in the UVOT range (1600-6000 Angstroms) are only half as bright as SN~2009dc, implying a smaller 56Ni yield. While not enough to strongly affect the bolometric flux, the early time mid-UV flux makes a significant contribution at early times. The strong spectral features in the mid-UV spectra of SNe 2009dc and 2012dn suggest a higher temperature and lower opacity to be the cause of the UV excess rather than a hot, smooth blackbody from shock interaction. Further work is needed to determine the ejecta and 56Ni masses of SNe 2011aa and 2012dn and fully explain their high UV luminosities.

astro-ph.HE

The Mass-Richness Relation of MaxBCG Clusters from Quasar Lensing Magnification using Variability

Accurate measurement of galaxy cluster masses is an essential component not only in studies of cluster physics, but also for probes of cosmology. However, different mass measurement techniques frequently yield discrepant results. The SDSS MaxBCG catalog's mass-richness relation has previously been constrained using weak lensing shear, Sunyaev-Zeldovich (SZ), and X-ray measurements. The mass normalization of the clusters as measured by weak lensing shear is >~25% higher than that measured using SZ and X-ray methods, a difference much larger than the stated measurement errors in the analyses. We constrain the mass-richness relation of the MaxBCG galaxy cluster catalog by measuring the gravitational lensing magnification of type I quasars in the background of the clusters. The magnification is determined using the quasars' variability and the correlation between quasars' variability amplitude and intrinsic luminosity. The mass-richness relation determined through magnification is in agreement with that measured using shear, confirming that the lensing strength of the clusters implies a high mass normalization, and that the discrepancy with other methods is not due to a shear-related systematic measurement error. We study the dependence of the measured mass normalization on the cluster halo orientation. As expected, line-of-sight clusters yield a higher normalization; however, this minority of haloes does not significantly bias the average mass-richness relation of the catalog.

astro-ph.CO