SearcharxivSearch

arXiv subjects

Joshua Wilde

Publications and source records attributed to Joshua Wilde.

3 recordsLinked to original sources

The impact of human expert visual inspection on the discovery of strong gravitational lenses

We investigate the ability of human 'expert' classifiers to identify strong gravitational lens candidates in Dark Energy Survey like imaging. We recruited a total of 55 people that completed more than 25$\%$ of the project. During the classification task, we present to the participants 1489 images. The sample contains a variety of data including lens simulations, real lenses, non-lens examples, and unlabeled data. We find that experts are extremely good at finding bright, well-resolved Einstein rings, whilst arcs with $g$-band signal-to-noise less than $\sim$25 or Einstein radii less than $\sim$1.2 times the seeing are rarely recovered. Very few non-lenses are scored highly. There is substantial variation in the performance of individual classifiers, but they do not appear to depend on the classifier's experience, confidence or academic position. These variations can be mitigated with a team of 6 or more independent classifiers. Our results give confidence that humans are a reliable pruning step for lens candidates, providing pure and quantifiably complete samples for follow-up studies.

astro-ph.GA

Detecting gravitational lenses using machine learning: exploring interpretability and sensitivity to rare lensing configurations

Forthcoming large imaging surveys such as Euclid and the Vera Rubin Observatory Legacy Survey of Space and Time are expected to find more than $10^5$ strong gravitational lens systems, including many rare and exotic populations such as compound lenses, but these $10^5$ systems will be interspersed among much larger catalogues of $\sim10^9$ galaxies. This volume of data is too much for visual inspection by volunteers alone to be feasible and gravitational lenses will only appear in a small fraction of these data which could cause a large amount of false positives. Machine learning is the obvious alternative but the algorithms' internal workings are not obviously interpretable, so their selection functions are opaque and it is not clear whether they would select against important rare populations. We design, build, and train several Convolutional Neural Networks (CNNs) to identify strong gravitational lenses using VIS, Y, J, and H bands of simulated data, with F1 scores between 0.83 and 0.91 on 100,000 test set images. We demonstrate for the first time that such CNNs do not select against compound lenses, obtaining recall scores as high as 76\% for compound arcs and 52\% for double rings. We verify this performance using Hubble Space Telescope (HST) and Hyper Suprime-Cam (HSC) data of all known compound lens systems. Finally, we explore for the first time the interpretability of these CNNs using Deep Dream, Guided Grad-CAM, and by exploring the kernels of the convolutional layers, to illuminate why CNNs succeed in compound lens selection.

astro-ph.GA

The Limits to Learning a Diffusion Model

This paper provides the first sample complexity lower bounds for the estimation of simple diffusion models, including the Bass model (used in modeling consumer adoption) and the SIR model (used in modeling epidemics). We show that one cannot hope to learn such models until quite late in the diffusion. Specifically, we show that the time required to collect a number of observations that exceeds our sample complexity lower bounds is large. For Bass models with low innovation rates, our results imply that one cannot hope to predict the eventual number of adopting customers until one is at least two-thirds of the way to the time at which the rate of new adopters is at its peak. In a similar vein, our results imply that in the case of an SIR model, one cannot hope to predict the eventual number of infections until one is approximately two-thirds of the way to the time at which the infection rate has peaked. This lower bound in estimation further translates into a lower bound in regret for decision-making in epidemic interventions. Our results formalize the challenge of accurate forecasting and highlight the importance of incorporating additional data sources. To this end, we analyze the benefit of a seroprevalence study in an epidemic, where we characterize the size of the study needed to improve SIR model estimation. Extensive empirical analyses on product adoption and epidemic data support our theoretical findings.

stat.ME