SearcharxivSearch

arXiv subjects

Chad Schafer

Publications and source records attributed to Chad Schafer.

12 recordsLinked to original sources

COWs and their Hybrids: A Statistical View of Custom Orthogonal Weights

A recurring challenge in high energy physics is inference of the signal component from a distribution for which observations are assumed to be a mixture of signal and background events. A standard assumption is that there exists information encoded in a discriminant variable that is effective at separating signal and background. This can be used to assign a signal weight to each event, with these weights used in subsequent analyses of one or more control variables of interest. The custom orthogonal weights (COWs) approach of Dembinski, et al.(2022), a generalization of the sPlot approach of Barlow (1987) and Pivk and Le Diberder (2005), is tailored to address this objective. The problem, and this method, present interesting and novel statistical issues. Here we formalize the assumptions needed and the statistical properties, while also considering extensions and alternative approaches.

stat.AP

Joint inference of multiplicative and additive systematics in galaxy density fluctuations and clustering measurements

Galaxy clustering measurements are a key probe of the matter density field in the Universe. With the era of precision cosmology upon us, surveys rely on precise measurements of the clustering signal for meaningful cosmological analysis. However, the presence of systematic contaminants can bias the observed galaxy number density, and thereby bias the galaxy two-point statistics. As the statistical uncertainties get smaller, correcting for these systematic contaminants becomes increasingly important for unbiased cosmological analysis. We present and validate a new method for understanding and mitigating both additive and multiplicative systematics in galaxy clustering measurements (two-point function) by joint inference of contaminants in the galaxy overdensity field (one-point function) using a maximum-likelihood estimator (MLE). We test this methodology with KiDS-like mock galaxy catalogs and synthetic systematic template maps. We estimate the cosmological impact of such mitigation by quantifying uncertainties and possible biases in the inferred relationship between the observed and the true galaxy clustering signal. Our method robustly corrects the clustering signal to the sub-percent level and reduces numerous additive and multiplicative systematics from $1.5 σ$ to less than $0.1σ$ for the scenarios we tested. In addition, we provide an empirical approach to identifying the functional form (additive, multiplicative, or other) by which specific systematics contaminate the galaxy number density. Even though this approach is tested and geared towards systematics contaminating the galaxy number density, the methods can be extended to systematics mitigation for other two-point correlation measurements.

astro-ph.CO

Reinterpreting Fundamental Plane Correlations with Machine Learning

This work explores the relationships between galaxy sizes and related observable galaxy properties in a large volume cosmological hydrodynamical simulation. The objectives of this work are to both develop a better understanding of the correlations between galaxy properties and the influence of environment on galaxy physics in order to build an improved model for the galaxy sizes, building off of the {\it fundamental plane}. With an accurate intrinsic galaxy size predictor, the residuals in the observed galaxy sizes can potentially be used for multiple cosmological applications, including making measurements of galaxy velocities in spectroscopic samples, estimating the rate of cosmic expansion, and constraining the uncertainties in the photometric redshifts of galaxies. Using projection pursuit regression, the model accurately predicts intrinsic galaxy sizes and have residuals which have limited correlation with galaxy properties. The model decreases the spatial correlation of galaxy size residuals by a factor of $\sim$ 5 at small scales compared to the baseline correlation when the mean size is used as a predictor.

astro-ph.GA

Algorithms and Statistical Models for Scientific Discovery in the Petabyte Era

The field of astronomy has arrived at a turning point in terms of size and complexity of both datasets and scientific collaboration. Commensurately, algorithms and statistical models have begun to adapt --- e.g., via the onset of artificial intelligence --- which itself presents new challenges and opportunities for growth. This white paper aims to offer guidance and ideas for how we can evolve our technical and collaborative frameworks to promote efficient algorithmic development and take advantage of opportunities for scientific discovery in the petabyte era. We discuss challenges for discovery in large and complex data sets; challenges and requirements for the next stage of development of statistical methodologies and algorithmic tool sets; how we might change our paradigms of collaboration and education; and the ethical implications of scientists' contributions to widely applicable algorithms and computational modeling. We start with six distinct recommendations that are supported by the commentary following them. This white paper is related to a larger corpus of effort that has taken place within and around the Petabytes to Science Workshops (https://petabytestoscience.github.io/).

astro-ph.IM

The Growing Importance of a Tech Savvy Astronomy and Astrophysics Workforce

Fundamental coding and software development skills are increasingly necessary for success in nearly every aspect of astronomical and astrophysical research as large surveys and high resolution simulations become the norm. However, professional training in these skills is inaccessible or impractical for many members of our community. Students and professionals alike have been expected to acquire these skills on their own, apart from formal classroom curriculum or on-the-job training. Despite the recognized importance of these skills, there is little opportunity to develop them - even for interested researchers. To ensure a workforce capable of taking advantage of the computational resources and the large volumes of data coming in the next decade, we must identify and support ways to make software development training widely accessible to community members, regardless of affiliation or career level. To develop and sustain a technology capable astronomical and astrophysical workforce, we recommend that agencies make funding and other resources available in order to encourage, support and, in some cases, require progress on necessary training, infrastructure and policies. In this white paper, we focus on recommendations for how funding agencies can lead in the promotion of activities to support the astronomy and astrophysical workforce in the 2020s.

astro-ph.IM

Realizing the potential of astrostatistics and astroinformatics

This Astro2020 State of the Profession Consideration White Paper highlights the growth of astrostatistics and astroinformatics in astronomy, identifies key issues hampering the maturation of these new subfields, and makes recommendations for structural improvements at different levels that, if acted upon, will make significant positive impacts across astronomy.

astro-ph.IM

Better support for collaborations preparing for large-scale projects: the case study of the LSST Science Collaborations Astro2020 APC White Paper

Through the lens of the LSST Science Collaborations' experience, this paper advocates for new and improved ways to fund large, complex collaborations at the interface of data science and astrophysics as they work in preparation for and on peta-scale, complex surveys, of which LSST is a prime example. We advocate for the establishment of programs to support both research and infrastructure development that enables innovative collaborative research on such scales.

astro-ph.IM

Astro2020 APC White Paper: Elevating the Role of Software as a Product of the Research Enterprise

Software is a critical part of modern research, and yet there are insufficient mechanisms in the scholarly ecosystem to acknowledge, cite, and measure the impact of research software. The majority of academic fields rely on a one-dimensional credit model whereby academic articles (and their associated citations) are the dominant factor in the success of a researcher's career. In the petabyte era of astronomical science, citing software and measuring its impact enables academia to retain and reward researchers that make significant software contributions. These highly skilled researchers must be retained to maximize the scientific return from petabyte-scale datasets. Evolving beyond the one-dimensional credit model requires overcoming several key challenges, including the current scholarly ecosystem and scientific culture issues. This white paper will present these challenges and suggest practical solutions for elevating the role of software as a product of the research enterprise.

astro-ph.IM

A Flexible Pipeline for Prediction of Tropical Cyclone Paths

Hurricanes and, more generally, tropical cyclones (TCs) are rare, complex natural phenomena of both scientific and public interest. The importance of understanding TCs in a changing climate has increased as recent TCs have had devastating impacts on human lives and communities. Moreover, good prediction and understanding about the complex nature of TCs can mitigate some of these human and property losses. Though TCs have been studied from many different angles, more work is needed from a statistical approach of providing prediction regions. The current state-of-the-art in TC prediction bands comes from the National Hurricane Center of the National Oceanographic and Atmospheric Administration (NOAA), whose proprietary model provides "cones of uncertainty" for TCs through an analysis of historical forecast errors. The contribution of this paper is twofold. We introduce a new pipeline that encourages transparent and adaptable prediction band development by streamlining cyclone track simulation and prediction band generation. We also provide updates to existing models and novel statistical methodologies in both areas of the pipeline, respectively.

stat.AP

A Preferential Attachment Model for the Stellar Initial Mass Function

Accurate specification of a likelihood function is becoming increasingly difficult in many inference problems in astronomy. As sample sizes resulting from astronomical surveys continue to grow, deficiencies in the likelihood function lead to larger biases in key parameter estimates. These deficiencies result from the oversimplification of the physical processes that generated the data, and from the failure to account for observational limitations. Unfortunately, realistic models often do not yield an analytical form for the likelihood. The estimation of a stellar initial mass function (IMF) is an important example. The stellar IMF is the mass distribution of stars initially formed in a given cluster of stars, a population which is not directly observable due to stellar evolution and other disruptions and observational limitations of the cluster. There are several difficulties with specifying a likelihood in this setting since the physical processes and observational challenges result in measurable masses that cannot legitimately be considered independent draws from an IMF. This work improves inference of the IMF by using an approximate Bayesian computation approach that both accounts for observational and astrophysical effects and incorporates a physically-motivated model for star cluster formation. The methodology is illustrated via a simulation study, demonstrating that the proposed approach can recover the true posterior in realistic situations, and applied to observations from astrophysical simulation data.

astro-ph.IM

Maximizing Science in the Era of LSST: A Community-Based Study of Needed US Capabilities

The Large Synoptic Survey Telescope (LSST) will be a discovery machine for the astronomy and physics communities, revealing astrophysical phenomena from the Solar System to the outer reaches of the observable Universe. While many discoveries will be made using LSST data alone, taking full scientific advantage of LSST will require ground-based optical-infrared (OIR) supporting capabilities, e.g., observing time on telescopes, instrumentation, computing resources, and other infrastructure. This community-based study identifies, from a science-driven perspective, capabilities that are needed to maximize LSST science. Expanding on the initial steps taken in the 2015 OIR System Report, the study takes a detailed, quantitative look at the capabilities needed to accomplish six representative LSST-enabled science programs that connect closely with scientific priorities from the 2010 decadal surveys. The study prioritizes the resources needed to accomplish the science programs and highlights ways that existing, planned, and future resources could be positioned to accomplish the science goals.

astro-ph.IM

Likelihood-Free Cosmological Inference with Type Ia Supernovae: Approximate Bayesian Computation for a Complete Treatment of Uncertainty

Cosmological inference becomes increasingly difficult when complex data-generating processes cannot be modeled by simple probability distributions. With the ever-increasing size of data sets in cosmology, there is increasing burden placed on adequate modeling; systematic errors in the model will dominate where previously these were swamped by statistical errors. For example, Gaussian distributions are an insufficient representation for errors in quantities like photometric redshifts. Likewise, it can be difficult to quantify analytically the distribution of errors that are introduced in complex fitting codes. Without a simple form for these distributions, it becomes difficult to accurately construct a likelihood function for the data as a function of parameters of interest. Approximate Bayesian computation (ABC) provides a means of probing the posterior distribution when direct calculation of a sufficiently accurate likelihood is intractable. ABC allows one to bypass direct calculation of the likelihood but instead relies upon the ability to simulate the forward process that generated the data. These simulations can naturally incorporate priors placed on nuisance parameters, and hence these can be marginalized in a natural way. We present and discuss ABC methods in the context of supernova cosmology using data from the SDSS-II Supernova Survey. Assuming a flat cosmology and constant dark energy equation of state we demonstrate that ABC can recover an accurate posterior distribution. Finally we show that ABC can still produce an accurate posterior distribution when we contaminate the sample with Type IIP supernovae.

astro-ph.CO