SearcharxivSearch

arXiv subjects

Tamas Budavari

Publications and source records attributed to Tamas Budavari.

At least 19 recordsLinked to original sources

Bayesian functional data analysis in astronomy

Cosmic demographics -- the statistical study of populations of astrophysical objects -- has long relied on *multivariate statistics*, providing methods for analyzing data comprising fixed-length vectors of properties of objects, as might be compiled in a tabular astronomical catalog (say, with sky coordinates, and brightness measurements in a fixed number of spectral passbands). But beginning with the emergence of automated digital sky surveys, ca. ~2000, astronomers began producing large collections of data with more complex structure: light curves (brightness time series) and spectra (brightness vs. wavelength). These comprise what statisticians call *functional data* -- measurements of populations of functions. Upcoming automated sky surveys will soon provide astronomers with a flood of functional data. New methods are needed to accurately and optimally analyze large ensembles of light curves and spectra, accumulating information both along and across measured functions. Functional data analysis (FDA) provides tools for statistical modeling of functional data. Astronomical data presents several challenges for FDA methodology, e.g., sparse, irregular, and asynchronous sampling, and heteroscedastic measurement error. Bayesian FDA uses hierarchical Bayesian models for function populations, and is well suited to addressing these challenges. We provide an overview of astronomical functional data, and of some key Bayesian FDA modeling approaches, including functional mixed effects models, and stochastic process models. We briefly describe a Bayesian FDA framework combining FDA and machine learning methods to build low-dimensional parametric models for galaxy spectra.

astro-ph.IM

A flexible Expectation-Maximization framework for fast, scalable and high-fidelity multi-frame astronomical image deconvolution

We present a computationally efficient expectation-maximization framework for multi-frame image deconvolution and super-resolution. Our method is well adapted for processing large scale imaging data from modern astronomical surveys. Our Tensorflow implementation is flexible, benefits from advanced algorithmic solutions, and allows users to seamlessly leverage Graphical Processing Unit (GPU) acceleration, thus making it viable for use in modern astronomical software pipelines. The testbed for our method is a set of $4$K by $4$K Hyper Suprime-Cam exposures, which are closest in terms of quality to imaging data from the upcoming Rubin Observatory. The preliminary results are extremely promising: our method produces a high-fidelity non-parametric reconstruction of the night sky, from which we recover unprecedented details such as the shape of the spiral arms of galaxies, while also managing to deconvolve stars perfectly into essentially single pixels.

astro-ph.IM

Learning the Night Sky with Deep Generative Priors

Recovering sharper images from blurred observations, referred to as deconvolution, is an ill-posed problem where classical approaches often produce unsatisfactory results. In ground-based astronomy, combining multiple exposures to achieve images with higher signal-to-noise ratios is complicated by the variation of point-spread functions across exposures due to atmospheric effects. We develop an unsupervised multi-frame method for denoising, deblurring, and coadding images inspired by deep generative priors. We use a carefully chosen convolutional neural network architecture that combines information from multiple observations, regularizes the joint likelihood over these observations, and allows us to impose desired constraints, such as non-negativity of pixel values in the sharp, restored image. With an eye towards the Rubin Observatory, we analyze 4K by 4K Hyper Suprime-Cam exposures and obtain preliminary results which yield promising restored images and extracted source lists.

cs.CV

Wireless sensor network for in situ soil moisture monitoring

We discuss the history and lessons learned from a series of deployments of environmental sensors measuring soil parameters and CO2 fluxes over the last fifteen years, in an outdoor environment. We present the hardware and software architecture of our current Gen-3 system, and then discuss how we are simplifying the user facing part of the software, to make it easier and friendlier for the environmental scientist to be in full control of the system. Finally, we describe the current effort to build a large-scale Gen-4 sensing platform consisting of hundreds of nodes to track the environmental parameters for urban green spaces in Baltimore, Maryland.

eess.SY

Subband Image Reconstruction using Differential Chromatic Refraction

Refraction by the atmosphere causes the positions of sources to depend on the airmass through which an observation was taken. This shift is dependent on the underlying spectral energy of the source and the filter or bandpass through which it is observed. Wavelength-dependent refraction within a single passband is often referred to as differential chromatic refraction (DCR). With a new generation of astronomical surveys undertaking repeated observations of the same part of the sky over a range of different airmasses and parallactic angles, DCR should be a detectable and measurable astrometric signal. In this paper we introduce a novel procedure that takes this astrometric signal and uses it to infer the underlying spectral energy distribution of a source; we solve for multiple latent images at specific wavelengths via a generalized deconvolution procedure built on robust statistics. We demonstrate the utility of such an approach for estimating a partially deconvolved image, at higher spectral resolution than the input images, for surveys such as the Large Synoptic Survey Telescope (LSST).

astro-ph.IM

Probabilistic Cross-identification of Multiple Catalogs in Crowded Fields

Matching astronomical catalogs in crowded regions of the sky is challenging both statistically and computationally due to the many possible alternative associations. Budavári and Basu (2016) modeled the two-catalog situation as an Assignment Problem and used the famous Hungarian algorithm to solve it. Here we treat cross-identification of multiple catalogs by introducing a different approach based on integer linear programming. We first test this new method on problems with two catalogs and compare with the previous results. We then test the efficacy of the new approach on problems with three catalogs. The performance and scalability of the new approach is discussed in the context of large surveys.

astro-ph.IM

Scalable Streaming Tools for Analyzing $N$-body Simulations: Finding Halos and Investigating Excursion Sets in One Pass

Cosmological $N$-body simulations play a vital role in studying models for the evolution of the Universe. To compare to observations and make a scientific inference, statistic analysis on large simulation datasets, e.g., finding halos, obtaining multi-point correlation functions, is crucial. However, traditional in-memory methods for these tasks do not scale to the datasets that are forbiddingly large in modern simulations. Our prior paper proposes memory-efficient streaming algorithms that can find the largest halos in a simulation with up to $10^9$ particles on a small server or desktop. However, this approach fails when directly scaling to larger datasets. This paper presents a robust streaming tool that leverages state-of-the-art techniques on GPU boosting, sampling, and parallel I/O, to significantly improve performance and scalability. Our rigorous analysis of the sketch parameters improves the previous results from finding the centers of the $10^3$ largest halos to $\sim 10^4-10^5$, and reveals the trade-offs between memory, running time and number of halos. Our experiments show that our tool can scale to datasets with up to $\sim 10^{12}$ particles while using less than an hour of running time on a single GPU Nvidia GTX 1080.

astro-ph.IM

Robust Statistics for Image Deconvolution

We present a blind multiframe image-deconvolution method based on robust statistics. The usual shortcomings of iterative optimization of the likelihood function are alleviated by minimizing the M-scale of the residuals, which achieves more uniform convergence across the image. We focus on the deconvolution of astronomical images, which are among the most challenging due to their huge dynamic ranges and the frequent presence of large noise-dominated regions in the images. We show that high-quality image reconstruction is possible even in super-resolution and without the use of traditional regularization terms. Using a robust \r{ho}-function is straightforward to implement in a streaming setting and, hence our method is applicable to the large volumes of astronomy images. The power of our method is demonstrated on observations from the Sloan Digital Sky Survey (Stripe 82) and we briefly discuss the feasibility of a pipeline based on Graphical Processing Units for the next generation of telescope surveys.

astro-ph.IM

Probabilistic Cross-Identification of Galaxies with Realistic Clustering

Probabilistic cross-identification has been successfully applied to a number of problems in astronomy from matching simple point sources to associating stars with unknown proper motions and even radio observations with realistic morphology. Here we study the Bayes factor for clustered objects and focus in particular on galaxies to assess the effect of typical angular correlations. Numerical calculations provide the modified relationship, which (as expected) suppresses the evidence for the associations at the shortest separations where the 2-point auto-correlation function is large. Ultimately this means that the matching probability drops at somewhat shorter scales than in previous models.

astro-ph.GA

Faint Object Detection in Multi-Epoch Observations via Catalog Data Fusion

Observational astronomy in the time-domain era faces several new challenges. One of them is the efficient use of observations obtained at multiple epochs. The work presented here addresses faint object detection with multi-epoch data, and describes an incremental strategy for separating real objects from artifacts in ongoing surveys, in situations where the single-epoch data are summaries of the full image data, such as single-epoch catalogs of flux and direction estimates for candidate sources. The basic idea is to produce low-threshold single-epoch catalogs, and use a probabilistic approach to accumulate catalog information across epochs; this is in contrast to more conventional strategies based on co-added or stacked image data across all epochs. We adopt a Bayesian approach, addressing object detection by calculating the marginal likelihoods for hypotheses asserting there is no object, or one object, in a small image patch containing at most one cataloged source at each epoch. The object-present hypothesis interprets the sources in a patch at different epochs as arising from a genuine object; the no-object (noise) hypothesis interprets candidate sources as spurious, arising from noise peaks. We study the detection probability for constant-flux objects in a simplified Gaussian noise setting, comparing results based on single exposures and stacked exposures to results based on a series of single-epoch catalog summaries. Computing the detection probability based on catalog data amounts to generalized cross-matching: it is the product of a factor accounting for matching of the estimated fluxes of candidate sources, and a factor accounting for matching of their estimated directions. We find that probabilistic fusion of multi-epoch catalog information can detect sources with only modest sacrifice in sensitivity and selectivity compared to stacking.

astro-ph.IM

Probabilistic record linkage in astronomy: Directional cross-identification and beyond

Modern astronomy increasingly relies upon systematic surveys, whose dedicated telescopes continuously observe the sky across varied wavelength ranges of the electromagnetic spectrum; some surveys also observe non-electromagnetic "messengers," such as high-energy particles or gravitational waves. Stars and galaxies look different through the eyes of different instruments, and their independent measurements have to be carefully combined to provide a complete, sound picture of the multicolor and eventful universe. The association of an object's independent detections is, however, a difficult problem scientifically, computationally, and statistically, raising varied challenges across diverse astronomical applications. The fundamental problem is finding records in survey databases with directions that match to within the direction uncertainties. Such astronomical versions of the record linkage problem are known by various terms in astronomy: cross-matching, cross-identification, and directional, positional, or spatio-temporal coincidence assessment. Astronomers have developed several statistical approaches for such problems, largely independently of related developments in other disciplines. Here we review emerging approaches that compute (Bayesian) probabilities for the hypotheses of interest: possible associations, or demographic properties of a cosmic population that depend on identifying associations. Many cross-identification tasks can be formulated within a hierarchical Bayesian partition model framework, with components that explicitly account for astrophysical effects (e.g., source brightness vs. wavelength, source motion, or source extent), selection effects, and measurement error. We survey recent developments, and highlight important open areas for future research.

astro-ph.IM

Probabilistic Cross-Identification in Crowded Fields as an Assignment Problem

One of the outstanding challenges of cross-identification is multiplicity: detections in crowded regions of the sky are often linked to more than one candidate associations of similar likelihoods. We map the resulting maximum likelihood partitioning to the fundamental assignment problem of discrete mathematics and efficiently solve the two-way catalog-level matching in the realm of combinatorial optimization using the so-called Hungarian algorithm. We introduce the method, demonstrate its performance in a mock universe where the true associations are known, and discuss the applicability of the new procedure to large surveys.

astro-ph.IM

Version 1 of the Hubble Source Catalog

The Hubble Source Catalog is designed to help optimize science from the Hubble Space Telescope by combining the tens of thousands of visit-based source lists in the Hubble Legacy Archive into a single master catalog. Version 1 of the Hubble Source Catalog includes WFPC2, ACS/WFC, WFC3/UVIS, and WFC3/IR photometric data generated using SExtractor software to produce the individual source lists. The catalog includes roughly 80 million detections of 30 million objects involving 112 different detector/filter combinations, and about 160 thousand HST exposures. Source lists from Data Release 8 of the Hubble Legacy Archive are matched using an algorithm developed by Budavari & Lubow (2012). The mean photometric accuracy for the catalog as a whole is better than 0.10 mag, with relative accuracy as good as 0.02 mag in certain circumstances (e.g., bright isolated stars). The relative astrometric residuals are typically within 10 mas, with a value for the mode (i.e., most common value) of 2.3 mas. The absolute astrometric accuracy is better than $\sim$0.1 arcsec for most sources, but can be much larger for a fraction of fields that could not be matched to the PanSTARRS, SDSS, or 2MASS reference systems. In this paper we describe the database design with emphasis on those aspects that enable the users to fully exploit the catalog while avoiding common misunderstandings and potential pitfalls. We provide usage examples to illustrate some of the science capabilities and data quality characteristics, and briefly discuss plans for future improvements to the Hubble Source Catalog.

astro-ph.IM

Clustering-based redshift estimation: method and application to data

We present a data-driven method to infer the redshift distribution of an arbitrary dataset based on spatial cross-correlation with a reference population and we apply it to various datasets across the electromagnetic spectrum to show its potential and limitations. Our approach advocates the use of clustering measurements on all available scales, in contrast to previous works focusing only on linear scales. We also show how its accuracy can be enhanced by optimally sampling a dataset within its photometric space rather than applying the estimator globally. We show that the ultimate goal of this technique is to characterize the mapping between the space of photometric observables and redshift space as this characterization then allows us to infer the clustering-redshift p.d.f. of a single galaxy. We apply this technique to estimate the redshift distributions of luminous red galaxies and emission line galaxies from the SDSS, infrared sources from WISE and radio sources from FIRST. We show that consistent redshift distributions are found using both quasars and absorber systems as reference populations. This technique brings valuable information on the third dimension of astronomical datasets. It is widely applicable to a large range of extra-galactic surveys.

astro-ph.CO

Objective Identification of Informative Wavelength Regions in Galaxy Spectra

Understanding the diversity in spectra is the key to determining the physical parameters of galaxies. The optical spectra of galaxies are highly convoluted with continuum and lines which are potentially sensitive to different physical parameters. Defining the wavelength regions of interest is therefore an important question. In this work, we identify informative wavelength regions in a single-burst stellar populations model by using the CUR Matrix Decomposition. Simulating the Lick/IDS spectrograph configuration, we recover the widely used Dn(4000), Hbeta, and HdeltaA to be most informative. Simulating the SDSS spectrograph configuration with a wavelength range 3450-8350 Angstrom and a model-limited spectral resolution of 3 Angstrom, the most informative regions are: first region-the 4000 Angstrom break and the Hdelta line; second region-the Fe-like indices; third region-the Hbeta line; fourth region-the G band and the Hgamma line. A Principal Component Analysis on the first region shows that the first eigenspectrum tells primarily the stellar age, the second eigenspectrum is related to the age-metallicity degeneracy, and the third eigenspectrum shows an anti-correlation between the strengths of the Balmer and the Ca K and H absorptions. The regions can be used to determine the stellar age and metallicity in early-type galaxies which have solar abundance ratios, no dust, and a single-burst star formation history. The region identification method can be applied to any set of spectra of the user's interest, so that we eliminate the need for a common, fixed-resolution index system. We discuss future directions in extending the current analysis to late-type galaxies.

astro-ph.CO

A Lyman Break Galaxy in the Epoch of Reionization from HST Grism Spectroscopy

We present observations of a luminous galaxy at redshift z=6.573 --- the end of the reioinization epoch --- which has been spectroscopically confirmed twice. The first spectroscopic confirmation comes from slitless HST ACS grism spectra from the PEARS survey (Probing Evolution And Reionization Spectroscopically), which show a dramatic continuum break in the spectrum at restframe 1216 A wavelength. The second confirmation is done with Keck + DEIMOS. The continuum is not clearly detected with ground-based spectra, but high wavelength resolution enables the Lyman alpha emission line profile to be determined. We compare the line profile to composite line profiles at redshift z=4.5. The Lyman alpha line profile shows no signature of a damping wing attenuation, confirming that the intergalactic gas is ionized at redshift z=6.57. Spectra of Lyman breaks at yet higher redshifts will be possible using comparably deep observations with IR-sensitive grisms, even at redshifts where Lyman alpha is too attenuated by the neutral IGM to be detectable using traditional spectroscopy from the ground.

astro-ph.CO

Catalog Matching with Astrometric Correction and its Application to the Hubble Legacy Archive

Object cross-identification in multiple observations is often complicated by the uncertainties in their astrometric calibration. Due to the lack of standard reference objects, an image with a small field of view can have significantly larger errors in its absolute positioning than the relative precision of the detected sources within. We present a new general solution for the relative astrometry that quickly refines the World Coordinate System of overlapping fields. The efficiency is obtained through the use of infinitesimal 3-D rotations on the celestial sphere, which do not involve trigonometric functions. They also enable an analytic solution to an important step in making the astrometric corrections. In cases with many overlapping images, the correct identification of detections that match together across different images is difficult to determine. We describe a new greedy Bayesian approach for selecting the best object matches across a large number of overlapping images. The methods are developed and demonstrated on the Hubble Legacy Archive, one of the most challenging data sets today. We describe a novel catalog compiled from many Hubble Space Telescope observations, where the detections are combined into a searchable collection of matches that link the individual detections. The matches provide descriptions of astronomical objects involving multiple wavelengths and epochs. High relative positional accuracy of objects is achieved across the Hubble images, often sub-pixel precision in the order of just a few milli-arcseconds. The result is a reliable set of high-quality associations that are publicly available online.

astro-ph.IM

Radio Continuum Surveys with Square Kilometre Array Pathfinders

In the lead-up to the Square Kilometre Array (SKA) project, several next-generation radio telescopes and upgrades are already being built around the world. These include APERTIF (The Netherlands), ASKAP (Australia), eMERLIN (UK), VLA (USA), e-EVN (based in Europe), LOFAR (The Netherlands), Meerkat (South Africa), and the Murchison Widefield Array (MWA). Each of these new instruments has different strengths, and coordination of surveys between them can help maximise the science from each of them. A radio continuum survey is being planned on each of them with the primary science objective of understanding the formation and evolution of galaxies over cosmic time, and the cosmological parameters and large-scale structures which drive it. In pursuit of this objective, the different teams are developing a variety of new techniques, and refining existing ones. Here we describe these projects, their science goals, and the technical challenges which are being addressed to maximise the science return.

astro-ph.CO