SearcharxivSearch

arXiv subjects

Hyungsuk Tak

Publications and source records attributed to Hyungsuk Tak.

At least 19 recordsLinked to original sources

Modeling Dependence Structures in Astronomical Multi-Band Time Series Data via Multi-Output Gaussian Processes

Modern astronomical time-domain surveys routinely collect multi-band light curves that provide complementary information about the physical processes governing source variability. Gaussian processes (GPs) provide a flexible probabilistic framework for modeling irregularly sampled and noisy time-series data. While considerable attention has been devoted to developing covariance kernels for individual time series, comparatively less attention has been paid to the statistical representation of dependence among multiple photometric bands. In this work, we present a unified statistical framework for modeling such dependence structures using multi-output GPs. Within this framework, we consider two complementary formulations. The covariance-based formulation specifies dependence directly through matrix-valued covariance functions and emphasizes the stochastic properties of the observed light curves, including covariance functions and power spectral densities. In contrast, the latent-process formulation represents the observed light curves as transformations of latent GPs and emphasizes the physical mechanisms generating the observed dependence. To illustrate these formulations, we develop covariance-based and latent-process multi-output damped random walk models and derive their corresponding spectral representations. We further demonstrate the practical implications of dependence-structure modeling through applications to multi-band active galactic nucleus variability and continuum reverberation mapping. Rather than advocating a universally preferred formulation, this work provides a principled basis for selecting dependence structures according to the scientific objectives and clarifies how this choice influences the statistical characterization and scientific interpretation of stochastic variability in astronomical sources.

astro-ph.IM

Six Maxims of Statistical Acumen for Astronomical Data Analysis

The production of complex astronomical data is accelerating, especially with newer telescopes producing ever more large-scale surveys. The increased quantity, complexity, and variety of astronomical data demand a parallel increase in skill and sophistication in developing, deciding, and deploying statistical methods. Understanding limitations and appreciating nuances in statistical and machine learning methods and the reasoning behind them is essential for improving data-analytic proficiency and acumen. Aiming to facilitate such improvement in astronomy, we delineate cautionary tales in statistics via six maxims, with examples drawn from the astronomical literature. Inspired by the significant quality improvement in business and manufacturing processes by the routine adoption of Six Sigma, we hope the routine reflection on these Six Maxims will improve the quality of both data analysis and scientific findings in astronomy.

astro-ph.IM

A Robust Bayesian Meta-Analysis for Estimating the Hubble Constant via Time Delay Cosmography

We propose a Bayesian meta-analysis to infer the current expansion rate of the Universe, called the Hubble constant ($H_0$), via time delay cosmography. Inputs of the meta-analysis are estimates of two properties for each pair of gravitationally lensed images; time delay and Fermat potential difference estimates with their standard errors. A meta-analysis can be appealing in practice because obtaining each estimate from even a single lens system involves substantial human efforts, and thus estimates are often separately obtained and published. Moreover, numerous estimates are expected to be available once the Rubin Observatory starts monitoring thousands of strong gravitational lens systems. This work focuses on combining these estimates from independent studies to infer $H_0$ in a robust manner. The robustness is crucial because currently up to eight lens systems are used to infer $H_0$, and thus any biased input can severely affect the resulting $H_0$ estimate. For this purpose, we adopt Student's $t$ error for the input estimates. We investigate properties of the resulting $H_0$ estimate via two simulation studies with realistic imaging data. It turns out that the meta-analysis can infer $H_0$ with sub-percent bias and about 1% level of coefficient of variation, even when 30% of inputs are manipulated to be outliers. We also apply the meta-analysis to three gravitationally lensed systems to obtain an $H_0$ estimate and compare it with existing estimates. An R package, h0, is publicly available for fitting the proposed meta-analysis.

stat.AP

Mapping the Growth of Supermassive Black Holes as a Function of Galaxy Stellar Mass and Redshift

The growth of supermassive black holes is strongly linked to their galaxies. It has been shown that the population mean black-hole accretion rate ($\overline{\mathrm{BHAR}}$) primarily correlates with the galaxy stellar mass ($M_\star$) and redshift for the general galaxy population. This work aims to provide the best measurements of $\overline{\mathrm{BHAR}}$ as a function of $M_\star$ and redshift over ranges of $10^{9.5}<M_\star<10^{12}~M_\odot$ and $z<4$. We compile an unprecedentedly large sample with eight thousand active galactic nuclei (AGNs) and 1.3 million normal galaxies from nine high-quality survey fields following a wedding-cake design. We further develop a semiparametric Bayesian method that can reasonably estimate $\overline{\mathrm{BHAR}}$ and the corresponding uncertainties, even for sparsely populated regions in the parameter space. $\overline{\mathrm{BHAR}}$ is constrained by X-ray surveys sampling the AGN accretion power and UV-to-infrared multi-wavelength surveys sampling the galaxy population. Our results can independently predict the X-ray luminosity function (XLF) from the galaxy stellar mass function (SMF), and the prediction is consistent with the observed XLF. We also try adding external constraints from the observed SMF and XLF. We further measure $\overline{\mathrm{BHAR}}$ for star-forming and quiescent galaxies and show that star-forming $\overline{\mathrm{BHAR}}$ is generally larger than or at least comparable to the quiescent $\overline{\mathrm{BHAR}}$.

astro-ph.GA

Repelling-Attracting Hamiltonian Monte Carlo

We propose a variant of Hamiltonian Monte Carlo (HMC), called the Repelling-Attracting Hamiltonian Monte Carlo (RAHMC), for sampling from multimodal distributions. The key idea that underpins RAHMC is a departure from the conservative dynamics of Hamiltonian systems, which form the basis of traditional HMC, and turning instead to the dissipative dynamics of conformal Hamiltonian systems. In particular, RAHMC involves two stages: a mode-repelling stage to encourage the sampler to move away from regions of high probability density; and, a mode-attracting stage, which facilitates the sampler to find and settle near alternative modes. We achieve this by introducing just one additional tuning parameter -- the coefficient of friction. The proposed method adapts to the geometry of the target distribution, e.g., modes and density ridges, and can generate proposals that cross low-probability barriers with little to no computational overhead in comparison to traditional HMC. Notably, RAHMC requires no additional information about the target distribution or memory of previously visited modes. We establish the theoretical basis for RAHMC, and we discuss repelling-attracting extensions to several variants of HMC in literature. Finally, we provide a tuning-free implementation via dual-averaging, and we demonstrate its effectiveness in sampling from, both, multimodal and unimodal distributions in high dimensions.

math.ST

TD-CARMA: Painless, accurate, and scalable estimates of gravitational-lens time delays with flexible CARMA processes

Cosmological parameters encoding our understanding of the expansion history of the Universe can be constrained by the accurate estimation of time delays arising in gravitationally lensed systems. We propose TD-CARMA, a Bayesian method to estimate cosmological time delays by modelling the observed and irregularly sampled light curves as realizations of a Continuous Auto-Regressive Moving Average (CARMA) process. Our model accounts for heteroskedastic measurement errors and microlensing, an additional source of independent extrinsic long-term variability in the source brightness. The semi-separable structure of the CARMA covariance matrix allows for fast and scalable likelihood computation using Gaussian Process modeling. We obtain a sample from the joint posterior distribution of the model parameters using a nested sampling approach. This allows for ``painless'' Bayesian Computation, dealing with the expected multi-modality of the posterior distribution in a straightforward manner and not requiring the specification of starting values or an initial guess for the time delay, unlike existing methods. In addition, the proposed sampling procedure automatically evaluates the Bayesian evidence, allowing us to perform principled Bayesian model selection. TD-CARMA is parsimonious, and typically includes no more than a dozen unknown parameters. We apply TD-CARMA to six doubly lensed quasars HS 2209+1914, SDSS J1001+5027, SDSS J1206+4332, SDSS J1515+1511, SDSS J1455+1447, SDSS J1349+1227, estimating their time delays as $-21.96 \pm 1.448$, $120.93 \pm 1.015$, $111.51 \pm 1.452$, $210.80 \pm 2.18$, $45.36 \pm 1.93$ and $432.05 \pm 1.950$ respectively. These estimates are consistent with those derived in the relevant literature, but are typically two to four times more precise.

astro-ph.IM

Practical Guidance for Bayesian Inference in Astronomy

In the last two decades, Bayesian inference has become commonplace in astronomy. At the same time, the choice of algorithms, terminology, notation, and interpretation of Bayesian inference varies from one sub-field of astronomy to the next, which can lead to confusion to both those learning and those familiar with Bayesian statistics. Moreover, the choice varies between the astronomy and statistics literature, too. In this paper, our goal is two-fold: (1) provide a reference that consolidates and clarifies terminology and notation across disciplines, and (2) outline practical guidance for Bayesian inference in astronomy. Highlighting both the astronomy and statistics literature, we cover topics such as notation, specification of the likelihood and prior distributions, inference using the posterior distribution, and posterior predictive checking. It is not our intention to introduce the entire field of Bayesian data analysis -- rather, we present a series of useful practices for astronomers who already have an understanding of the Bayesian "nuts and bolts" and wish to increase their expertise and extend their knowledge. Moreover, as the field of astrostatistics and astroinformatics continues to grow, we hope this paper will serve as both a helpful reference and as a jumping off point for deeper dives into the statistics and astrostatistics literature.

astro-ph.IM

Incorporating Measurement Error in Astronomical Object Classification

Most general-purpose classification methods, such as support-vector machine (SVM) and random forest (RF), fail to account for an unusual characteristic of astronomical data: known measurement error uncertainties. In astronomical data, this information is often given in the data but discarded because popular machine learning classifiers cannot incorporate it. We propose a simulation-based approach that incorporates heteroscedastic measurement error into existing classification method to better quantify uncertainty in classification. The proposed method first simulates perturbed realizations of the data from a Bayesian posterior predictive distribution of a Gaussian measurement error model. Then, a chosen classifier is fit to each simulation. The variation across the simulations naturally reflects the uncertainty propagated from the measurement errors in both labeled and unlabeled data sets. We demonstrate the use of this approach via two numerical studies. The first is a thorough simulation study applying the proposed procedure to SVM and RF, which are well-known hard and soft classifiers, respectively. The second study is a realistic classification problem of identifying high-$z$ $(2.9 \leq z \leq 5.1)$ quasar candidates from photometric data. The data are from merged catalogs of the Sloan Digital Sky Survey, the $Spitzer$ IRAC Equatorial Survey, and the $Spitzer$-HETDEX Exploratory Large-Area Survey. The proposed approach reveals that out of 11,847 high-$z$ quasar candidates identified by a random forest without incorporating measurement error, 3,146 are potential misclassifications with measurement error. Additionally, out of $1.85$ million objects not identified as high-$z$ quasars without measurement error, 936 can be considered new candidates with measurement error.

astro-ph.IM

Modeling Stochastic Variability in Multi-Band Time Series Data

In preparation for the era of the time-domain astronomy with upcoming large-scale surveys, we propose a state-space representation of a multivariate damped random walk process as a tool to analyze irregularly-spaced multi-filter light curves with heteroscedastic measurement errors. We adopt a computationally efficient and scalable Kalman-filtering approach to evaluate the likelihood function, leading to maximum $O(k^3n)$ complexity, where $k$ is the number of available bands and $n$ is the number of unique observation times across the $k$ bands. This is a significant computational advantage over a commonly used univariate Gaussian process that can stack up all multi-band light curves in one vector with maximum $O(k^3n^3)$ complexity. Using such efficient likelihood computation, we provide both maximum likelihood estimates and Bayesian posterior samples of the model parameters. Three numerical illustrations are presented; (i) analyzing simulated five-band light curves for a comparison with independent single-band fits; (ii) analyzing five-band light curves of a quasar obtained from the Sloan Digital Sky Survey (SDSS) Stripe~82 to estimate the short-term variability and timescale; (iii) analyzing gravitationally lensed $g$- and $r$-band light curves of Q0957+561 to infer the time delay. Two R packages, Rdrw and timedelay, are publicly available to fit the proposed models.

astro-ph.IM

Data transforming augmentation for heteroscedastic models

Data augmentation (DA) turns seemingly intractable computational problems into simple ones by augmenting latent missing data. In addition to computational simplicity, it is now well-established that DA equipped with a deterministic transformation can improve the convergence speed of iterative algorithms such as an EM algorithm or Gibbs sampler. In this article, we outline a framework for the transformation-based DA, which we call data transforming augmentation (DTA), allowing augmented data to be a deterministic function of latent and observed data, and unknown parameters. Under this framework, we investigate a novel DTA scheme that turns heteroscedastic models into homoscedastic ones to take advantage of simpler computations typically available in homoscedastic cases. Applying this DTA scheme to fitting linear mixed models, we demonstrate simpler computations and faster convergence rates of resulting iterative algorithms, compared with those under a non-transformation-based DA scheme. We also fit a Beta-Binomial model using the proposed DTA scheme, which enables sampling approximate marginal posterior distributions that are available only under homoscedasticity. An R package, Rdta, is publicly available at CRAN.

stat.ME

How proper are Bayesian models in the astronomical literature?

The well-known Bayes theorem assumes that a posterior distribution is a probability distribution. However, the posterior distribution may no longer be a probability distribution if an improper prior distribution (non-probability measure) such as an unbounded uniform prior is used. Improper priors are often used in the astronomical literature to reflect a lack of prior knowledge, but checking whether the resulting posterior is a probability distribution is sometimes neglected. It turns out that 23 articles out of 75 articles (30.7%) published online in two renowned astronomy journals (ApJ and MNRAS) between Jan 1, 2017 and Oct 15, 2017 make use of Bayesian analyses without rigorously establishing posterior propriety. A disturbing aspect is that a Gibbs-type Markov chain Monte Carlo (MCMC) method can produce a seemingly reasonable posterior sample even when the posterior is not a probability distribution (Hobert and Casella, 1996). In such cases, researchers may erroneously make probabilistic inferences without noticing that the MCMC sample is from a non-existing probability distribution. We review why checking posterior propriety is fundamental in Bayesian analyses, and discuss how to set up scientifically motivated proper priors.

astro-ph.IM

Robust and Accurate Inference via a Mixture of Gaussian and Student's t Errors

A Gaussian measurement error assumption, i.e., an assumption that the data are observed up to Gaussian noise, can bias any parameter estimation in the presence of outliers. A heavy tailed error assumption based on Student's t distribution helps reduce the bias. However, it may be less efficient in estimating parameters if the heavy tailed assumption is uniformly applied to all of the data when most of them are normally observed. We propose a mixture error assumption that selectively converts Gaussian errors into Student's t errors according to latent outlier indicators, leveraging the best of the Gaussian and Student's t errors; a parameter estimation can be not only robust but also accurate. Using simulated hospital profiling data and astronomical time series of brightness data, we demonstrate the potential for the proposed mixture error assumption to estimate parameters accurately in the presence of outliers.

stat.ME

Rapid Optical Variations Correlated with X-rays in the 2015 Second Outburst of V404 Cygni (GS 2023$+$338)

We present optical multi-colour photometry of V404 Cyg during the outburst from December, 2015 to January, 2016 together with the simultaneous X-ray data. This outburst occurred less than 6 months after the previous outburst in June-July, 2015. These two outbursts in 2015 were of a slow rise and rapid decay-type and showed large-amplitude ($\sim$2 mag) and short-term ($\sim$10 min-3 hours) optical variations even at low luminosity (0.01-0.1$L_{\rm Edd}$). We found correlated optical and X-ray variations in two $\sim$1 hour time intervals and performed Bayesian time delay estimations between them. In the previous version, the observation times of X-ray light curves were measured at the satellite and their system of times was Terrestrial Time (TT), while those of optical light curves were measured at the Earth and their system of times was Coordinated Universal Time (UTC). In this version, we have corrected the observation times and obtained a Bayesian estimate of an optical delay against the X-ray emission, which is $\sim$30 s, during those two intervals. In addition, the relationship between the optical and X-ray luminosity was $L_{\rm opt} \propto L_{\rm X}^{0.25-0.29}$ at that time. These features can be naturally explained by disc reprocessing.

astro-ph.HE

A Repelling-Attracting Metropolis Algorithm for Multimodality

Although the Metropolis algorithm is simple to implement, it often has difficulties exploring multimodal distributions. We propose the repelling-attracting Metropolis (RAM) algorithm that maintains the simple-to-implement nature of the Metropolis algorithm, but is more likely to jump between modes. The RAM algorithm is a Metropolis-Hastings algorithm with a proposal that consists of a downhill move in density that aims to make local modes repelling, followed by an uphill move in density that aims to make local modes attracting. The downhill move is achieved via a reciprocal Metropolis ratio so that the algorithm prefers downward movement. The uphill move does the opposite using the standard Metropolis ratio which prefers upward movement. This down-up movement in density increases the probability of a proposed move to a different mode. Because the acceptance probability of the proposal involves a ratio of intractable integrals, we introduce an auxiliary variable which creates a term in the acceptance probability that cancels with the intractable ratio. Using several examples, we demonstrate the potential for the RAM algorithm to explore a multimodal distribution more efficiently than a Metropolis algorithm and with less tuning than is commonly required by tempering-based methods.

stat.ME

Bayesian Estimates of Astronomical Time Delays between Gravitationally Lensed Stochastic Light Curves

The gravitational field of a galaxy can act as a lens and deflect the light emitted by a more distant object such as a quasar. Strong gravitational lensing causes multiple images of the same quasar to appear in the sky. Since the light in each gravitationally lensed image traverses a different path length from the quasar to the Earth, fluctuations in the source brightness are observed in the several images at different times. The time delay between these fluctuations can be used to constrain cosmological parameters and can be inferred from the time series of brightness data or light curves of each image. To estimate the time delay, we construct a model based on a state-space representation for irregularly observed time series generated by a latent continuous-time Ornstein-Uhlenbeck process. We account for microlensing, an additional source of independent long-term extrinsic variability, via a polynomial regression. Our Bayesian strategy adopts a Metropolis-Hastings within Gibbs sampler. We improve the sampler by using an ancillarity-sufficiency interweaving strategy and adaptive Markov chain Monte Carlo. We introduce a profile likelihood of the time delay as an approximation of its marginal posterior distribution. The Bayesian and profile likelihood approaches complement each other, producing almost identical results; the Bayesian method is more principled but the profile likelihood is simpler to implement. We demonstrate our estimation strategy using simulated data of doubly- and quadruply-lensed quasars, and observed data from quasars Q0957+561 and J1029+2623.

astro-ph.IM

Frequency Coverage Properties of a Uniform Shrinkage Prior Distribution

A uniform shrinkage prior (USP) distribution on the unknown variance component of a random-effects model is known to produce good frequency properties. The USP has a parameter that determines the shape of its density function, but it has been neglected whether the USP can maintain such good frequency properties regardless of the choice for the shape parameter. We investigate which choice for the shape parameter of the USP produces Bayesian interval estimates of random effects that meet their nominal confidence levels better than several existent choices in the literature. Using univariate and multivariate Gaussian hierarchical models, we empirically show that the USP can achieve its best frequency properties when its shape parameter makes the USP behave similarly to an improper flat prior distribution on the unknown variance component.

stat.ME

Rgbp: An R Package for Gaussian, Poisson, and Binomial Random Effects Models with Frequency Coverage Evaluations

Rgbp is an R package that provides estimates and verifiable confidence intervals for random effects in two-level conjugate hierarchical models for overdispersed Gaussian, Poisson, and Binomial data. Rgbp models aggregate data from k independent groups summarized by observed sufficient statistics for each random effect, such as sample means, possibly with covariates. Rgbp uses approximate Bayesian machinery with unique improper priors for the hyper-parameters, which leads to good repeated sampling coverage properties for random effects. A special feature of Rgbp is an option that generates synthetic data sets to check whether the interval estimates for random effects actually meet the nominal confidence levels. Additionally, Rgbp provides inference statistics for the hyper-parameters, e.g., regression coefficients.

stat.ME

Data-dependent Posterior Propriety of Bayesian Beta-Binomial-Logit Model

A Beta-Binomial-Logit model is a Beta-Binomial model with covariate information incorporated via a logistic regression. Posterior propriety of a Bayesian Beta-Binomial-Logit model can be data-dependent for improper hyper-prior distributions. Various researchers in the literature have unknowingly used improper posterior distributions or have given incorrect statements about posterior propriety because checking posterior propriety can be challenging due to the complicated functional form of a Beta-Binomial-Logit model. We derive data-dependent necessary and sufficient conditions for posterior propriety within a class of hyper-prior distributions that encompass those used in previous studies.

math.ST