Searcharxiv⌕ Search

arXiv subjects

Michael D. Schneider

Publications and source records attributed to Michael D. Schneider.

At least 19 recordsLinked to original sources

Markov Chain Monte Carlo for Bayesian Parametric Galaxy Modeling in LSST

We apply Markov Chain Monte Carlo (MCMC) to the problem of parametric galaxy modeling, estimating posterior distributions of galaxy properties such as ellipticity and brightness for more than 100,000 coadded images of galaxies taken from DC2, a simulated telescope survey resembling the ongoing Rubin Observatory Legacy Survey of Space and Time (LSST). This analysis focuses only on truly unblended galaxies detected as single objects. We use a physically informed prior, apply selection corrections to the likelihood, and systematically study the bias and calibration of our posteriors. The resulting posterior samples support rigorous probabilistic inference of galaxy model parameters and their uncertainties, even for low signal-to-noise galaxies that are often excluded from cosmological analyses. We implement the probabilistic modeling and MCMC inference using the JIF (Joint Image Framework) package, which we make freely available.

astro-ph.IM↗

Control-affine Schrödinger Bridge and Generalized Bohm Potential

The control-affine Schrödinger bridge concerns with a stochastic optimal control problem. Its solution is a controlled evolution of joint state probability density subject to a control-affine Itô diffusion with a given deadline connecting a given pair of initial and terminal densities. In this work, we recast the necessary conditions of optimality for the control-affine Schrödinger bridge problem as a two point boundary value problem for a quantum mechanical Schrödinger PDE with complex potential. This complex-valued potential is a generalization of the real-valued Bohm potential in quantum mechanics. Our derived potential is akin to the optical potential in nuclear physics where the real part of the potential encodes elastic scattering (transmission of wave function), and the imaginary part encodes inelastic scattering (absorption of wave function). The key takeaway is that the process noise that drives the evolution of probability densities induces an absorbing medium in the evolution of wave function. These results make new connections between control theory and non-equilibrium statistical mechanics through the lens of quantum mechanics.

math.OC↗

Identifiability and Sensitivity Analysis of Kriging Weights for the Matern Kernel

Gaussian process (GP) models are effective non-linear models for numerous scientific applications. However, computation of their hyperparameters can be difficult when there is a large number of training observations (n) due to the O(n^3) cost of evaluating the likelihood function. Furthermore, non-identifiable hyperparameter values can induce difficulty in parameter estimation. Because of this, maximum likelihood estimation or Bayesian calibration is sometimes omitted and the hyperparameters are estimated with prediction-based methods such as a grid search using cross validation. Kriging, or prediction using a Gaussian process model, amounts to a weighted mean of the data, where training data close to the prediction location as determined by the form and hyperparameters of the kernel matrix are more highly weighted. Our analysis focuses on examination of the commonly utilized Matern covariance function, of which the radial basis function (RBF) kernel function is the infinity limit of the smoothness parameter. We first perform a collinearity analysis to motivate identifiability issues between the parameters of the Matern covariance function. We also demonstrate which of its parameters can be estimated using only the predictions. Considering the kriging weights for a fixed training data and prediction location as a function of the hyperparameters, we evaluate their sensitivities - as well as those of the predicted variance - with respect to said hyperparameters. We demonstrate the smoothness parameter nu is the most sensitive parameter in determining the kriging weights, particularly when the nugget parameter is small, indicating this is the most important parameter to estimate. Finally, we demonstrate the impact of our conclusions on performance and accuracy in a classification problem using a latent Gaussian process model with the hyperparameters selected via a grid search.

stat.CO↗

Correspondence of NNGP Kernel and the Matern Kernel

Kernels representing limiting cases of neural network architectures have recently gained popularity. However, the application and performance of these new kernels compared to existing options, such as the Matern kernel, is not well studied. We take a practical approach to explore the neural network Gaussian process (NNGP) kernel and its application to data in Gaussian process regression. We first demonstrate the necessity of normalization to produce valid NNGP kernels and explore related numerical challenges. We further demonstrate that the predictions from this model are quite inflexible, and therefore do not vary much over the valid hyperparameter sets. We then demonstrate a surprising result that the predictions given from the NNGP kernel correspond closely to those given by the Matern kernel under specific circumstances, which suggests a deep similarity between overparameterized deep neural networks and the Matern kernel. Finally, we demonstrate the performance of the NNGP kernel as compared to the Matern kernel on three benchmark data cases, and we conclude that for its flexibility and practical performance, the Matern kernel is preferred to the novel NNGP in practical applications.

stat.ML↗

A Scalable Gaussian Process Approach to Shear Mapping with MuyGPs

Analysis of cosmic shear is an integral part of understanding structure growth across cosmic time, which in-turn provides us with information about the nature of dark energy. Conventional methods generate \emph{shear maps} from which we can infer the matter distribution in the universe. Current methods (e.g., Kaiser-Squires inversion) for generating these maps, however, are tricky to implement and can introduce bias. Recent alternatives construct a spatial process prior for the lensing potential, which allows for inference of the convergence and shear parameters given lensing shear measurements. Realizing these spatial processes, however, scales cubically in the number of observations - an unacceptable expense as near-term surveys expect billions of correlated measurements. Therefore, we present a linearly-scaling shear map construction alternative using a scalable Gaussian Process (GP) prior called MuyGPs. MuyGPs avoids cubic scaling by conditioning interpolation on only nearest-neighbors and fits hyperparameters using batched leave-one-out cross validation. We use a suite of ray-tracing results from N-body simulations to demonstrate that our method can accurately interpolate shear maps, as well as recover the two-point and higher order correlations. We also show that we can perform these operations at the scale of billions of galaxies on high performance computing platforms.

astro-ph.CO↗

GREAT3 results I: systematic errors in shear estimation and the impact of real galaxy morphology

We present first results from the third GRavitational lEnsing Accuracy Testing (GREAT3) challenge, the third in a sequence of challenges for testing methods of inferring weak gravitational lensing shear distortions from simulated galaxy images. GREAT3 was divided into experiments to test three specific questions, and included simulated space- and ground-based data with constant or cosmologically-varying shear fields. The simplest (control) experiment included parametric galaxies with a realistic distribution of signal-to-noise, size, and ellipticity, and a complex point spread function (PSF). The other experiments tested the additional impact of realistic galaxy morphology, multiple exposure imaging, and the uncertainty about a spatially-varying PSF; the last two questions will be explored in Paper II. The 24 participating teams competed to estimate lensing shears to within systematic error tolerances for upcoming Stage-IV dark energy surveys, making 1525 submissions overall. GREAT3 saw considerable variety and innovation in the types of methods applied. Several teams now meet or exceed the targets in many of the tests conducted (to within the statistical errors). We conclude that the presence of realistic galaxy morphology in simulations changes shear calibration biases by $\sim 1$ per cent for a wide range of methods. Other effects such as truncation biases due to finite galaxy postage stamps, and the impact of galaxy type as measured by the Sérsic index, are quantified for the first time. Our results generalize previous studies regarding sensitivities to galaxy size and signal-to-noise, and to PSF properties such as seeing and defocus. Almost all methods' results support the simple model in which additive shear biases depend linearly on PSF ellipticity.

astro-ph.CO↗

Optimal Control From Inverse Scattering via Single-Sided Focusing

We describe an algorithm to solve Bellman optimization that replaces a sum over paths determining the optimal cost-to-go by an analytic method localized in state space. Our approach follows from the established relation between stochastic control problems in the class of linear Markov decision processes and quantum inverse scattering. We introduce a practical online computational method to solve for a potential function that informs optimal agent actions. This approach suggests that optimal control problems, including those with many degrees of freedom, can be solved with parallel computations.

math.OC↗

Gaussian Process Classification for Galaxy Blend Identification in LSST

A significant fraction of observed galaxies in the Rubin Observatory Legacy Survey of Space and Time (LSST) will overlap at least one other galaxy along the same line of sight, in a so-called "blend." The current standard method of assessing blend likelihood in LSST images relies on counting up the number of intensity peaks in the smoothed image of a blend candidate, but the reliability of this procedure has not yet been comprehensively studied. Here we construct a realistic distribution of blended and unblended galaxies through high-fidelity simulations of LSST-like images, and from this we examine the blend classification accuracy of the standard peak-finding method. Furthermore, we develop a novel Gaussian process blend classifier model, and show that this classifier is competitive with both the peak-finding method as well as with a convolutional neural network model. Finally, whereas the peak-finding method does not naturally assign probabilities to its classification estimates, the Gaussian process model does, and we show that the Gaussian process classification probabilities are generally reliable.

astro-ph.IM↗

Rare Events via Cross-Entropy Population Monte Carlo

We present a Cross-Entropy based population Monte Carlo algorithm. This methods stands apart from previous work in that we are not optimizing a mixture distribution. Instead, we leverage deterministic mixture weights and optimize the distributions individually through a reinterpretation of the typical derivation of the cross-entropy method. Demonstrations on numerical examples show that the algorithm can outperform existing resampling population Monte Carlo methods, especially for higher-dimensional problems.

stat.CO↗

Star-Galaxy Image Separation with Computationally Efficient Gaussian Process Classification

We introduce a novel method for discerning optical telescope images of stars from those of galaxies using Gaussian processes (GPs). Although applications of GPs often struggle in high-dimensional data modalities such as optical image classification, we show that a low-dimensional embedding of images into a metric space defined by the principal components of the data suffices to produce high-quality predictions from real large-scale survey data. We develop a novel method of GP classification hyperparameter training that scales approximately linearly in the number of image observations, which allows for application of GP models to large-size Hyper Suprime-Cam (HSC) Subaru Strategic Program data. In our experiments we evaluate the performance of a principal component analysis (PCA) embedded GP predictive model against other machine learning algorithms including a convolutional neural network and an image photometric morphology discriminator. Our analysis shows that our methods compare favorably with current methods in optical image classification while producing posterior distributions from the GP regression that can be used to quantify object classification uncertainty. We further describe how classification uncertainty can be used to efficiently parse large-scale survey imaging data to produce high-confidence object catalogs.

astro-ph.IM↗

A New Blind Asteroid Detection Scheme

As astronomical photometric surveys continue to tile the sky repeatedly, the potential to pushdetection thresholds to fainter limits increases; however, traditional digital-tracking methods cannotachieve this efficiently beyond time scales where motion is approximately linear. In this paper weprototype an optimal detection scheme that samples under a user defined prior on a parameterizationof the motion space, maps these sampled trajectories to the data space, and computes an optimalsignal-matched filter for computing the signal to noise ratio of trial trajectories. We demonstrate thecapability of this method on a small test data set from the Dark Energy Camera. We recover themajority of asteroids expected to appear and also discover hundreds of new asteroids with only a fewhours of observations. We conclude by exploring the potential for extending this scheme to larger datasets that cover larger areas of the sky over longer time baselines.

astro-ph.EP↗

The Impact of Tomographic Redshift Bin Width Errors on Cosmological Probes

Systematic errors in the galaxy redshift distribution $n(z)$ can propagate to systematic errors in the derived cosmology. We characterize how the degenerate effects in tomographic bin widths and galaxy bias impart systematic errors on cosmology inference using observational data from the Deep Lens Survey. For this we use a combination of galaxy clustering and galaxy-galaxy lensing. We present two end-to-end analyses from the catalogue level to parameter estimation. We produce an initial cosmological inference using fiducial tomographic redshift bins derived from photometric redshifts, then compare this with a result where the redshift bins are empirically corrected using a set of spectroscopic redshifts. We find that the derived parameter $S_8 \equiv σ_8 (Ω_m/.3)^{1/2}$ goes from $.841^{+0.062}_{-.061}$ to $.739^{+.054}_{-.050}$ upon correcting the n(z) errors in the second method.

astro-ph.CO↗

Bayesian Fusion of Data Partitioned Particle Estimates

We present a Bayesian data fusion method to approximate a posterior distribution from an ensemble of particle estimates that only have access to subsets of the data. Our approach relies on approximate probabilistic inference of model parameters through Monte Carlo methods, followed by an update and resample scheme related to multiple importance sampling to combine information from the initial estimates. We show the method is convergent in the particle limit and directly suited to application on multi-sensor data fusion problems by demonstrating efficacy on a multi-sensor Keplerian orbit determination problem and a bearings-only tracking problem.

stat.CO↗

Star-Galaxy Separation via Gaussian Processes with Model Reduction

Modern cosmological surveys such as the Hyper Suprime-Cam (HSC) survey produce a huge volume of low-resolution images of both distant galaxies and dim stars in our own galaxy. Being able to automatically classify these images is a long-standing problem in astronomy and critical to a number of different scientific analyses. Recently, the challenge of "star-galaxy" classification has been approached with Deep Neural Networks (DNNs), which are good at learning complex nonlinear embeddings. However, DNNs are known to overconfidently extrapolate on unseen data and require a large volume of training images that accurately capture the data distribution to be considered reliable. Gaussian Processes (GPs), which infer posterior distributions over functions and naturally quantify uncertainty, haven't been a tool of choice for this task mainly because popular kernels exhibit limited expressivity on complex and high-dimensional data. In this paper, we present a novel approach to the star-galaxy separation problem that uses GPs and reap their benefits while solving many of the issues traditionally affecting them for classification of high-dimensional celestial image data. After an initial filtering of the raw data of star and galaxy image cutouts, we first reduce the dimensionality of the input images by using a Principal Components Analysis (PCA) before applying GPs using a simple Radial Basis Function (RBF) kernel on the reduced data. Using this method, we greatly improve the accuracy of the classification over a basic application of GPs while improving the computational efficiency and scalability of the method.

astro-ph.IM↗

A Reanalysis of Public Galactic Bulge Gravitational Microlensing Events from OGLE-III and IV

Modern surveys of gravitational microlensing events have progressed to detecting thousands per year. Surveys are capable of probing Galactic structure, stellar evolution, lens populations, black hole physics, and the nature of dark matter. One of the key avenues for doing this is studying the microlensing Einstein radius crossing time distribution ($t_E$). However, systematics in individual light curves as well as over-simplistic modeling can lead to biased results. To address this, we developed a model to simultaneously handle the microlensing parallax due to Earth's motion, systematic instrumental effects, and unlensed stellar variability with a Gaussian Process model. We used light curves for nearly 10,000 OGLE-III and IV Milky Way bulge microlensing events and fit each with our model. We also developed a forward model approach to infer the timescale distribution by forward modeling from the data rather than using point estimates from individual events. We find that modeling the variability in the baseline removes a source of significant bias in individual events, and previous analyses over-estimated the number of long timescale ($t_E>100$ days) events due to their over simplistic models ignoring parallax effects and stellar variability. We use our fits to identify hundreds of events that are likely black holes.

astro-ph.GA↗

Quantum Machine Learning using Gaussian Processes with Performant Quantum Kernels

Quantum computers have the opportunity to be transformative for a variety of computational tasks. Recently, there have been proposals to use the unsimulatably of large quantum devices to perform regression, classification, and other machine learning tasks with quantum advantage by using kernel methods. While unsimulatably is a necessary condition for quantum advantage in machine learning, it is not sufficient, as not all kernels are equally effective. Here, we study the use of quantum computers to perform the machine learning tasks of one- and multi-dimensional regression, as well as reinforcement learning, using Gaussian Processes. By using approximations of performant classical kernels enhanced with extra quantum resources, we demonstrate that quantum devices, both in simulation and on hardware, can perform machine learning tasks at least as well as, and many times better than, the classical inspiration. Our informed kernel design demonstrates a path towards effectively utilizing quantum devices for machine learning tasks.

quant-ph↗

Reinforcement Learning via Gaussian Processes with Neural Network Dual Kernels

While deep neural networks (DNNs) and Gaussian Processes (GPs) are both popularly utilized to solve problems in reinforcement learning, both approaches feature undesirable drawbacks for challenging problems. DNNs learn complex nonlinear embeddings, but do not naturally quantify uncertainty and are often data-inefficient to train. GPs infer posterior distributions over functions, but popular kernels exhibit limited expressivity on complex and high-dimensional data. Fortunately, recently discovered conjugate and neural tangent kernel functions encode the behavior of overparameterized neural networks in the kernel domain. We demonstrate that these kernels can be efficiently applied to regression and reinforcement learning problems by analyzing a baseline case study. We apply GPs with neural network dual kernels to solve reinforcement learning tasks for the first time. We demonstrate, using the well-understood mountain-car problem, that GPs empowered with dual kernels perform at least as well as those using the conventional radial basis function kernel. We conjecture that by inheriting the probabilistic rigor of GPs and the powerful embedding properties of DNNs, GPs using NN dual kernels will empower future reinforcement learning models on difficult domains.

cs.LG↗

Probabilistic Cosmological Mass Mapping from Weak Lensing Shear

We infer gravitational lensing shear and convergence fields from galaxy ellipticity catalogs under a spatial process prior for the lensing potential. We demonstrate the performance of our algorithm with simulated Gaussian-distributed cosmological lensing shear maps and a reconstruction of the mass distribution of the merging galaxy cluster Abell 781 using galaxy ellipticities measured with the Deep Lens Survey. Given interim posterior samples of lensing shear or convergence fields on the sky, we describe an algorithm to infer cosmological parameters via lens field marginalization. In the most general formulation of our algorithm we make no assumptions about weak shear or Gaussian distributed shape noise or shears. Because we require solutions and matrix determinants of a linear system of dimension that scales with the number of galaxies, we expect our algorithm to require parallel high-performance computing resources for application to ongoing wide field lensing surveys.

astro-ph.CO↗