SearcharxivSearch

arXiv subjects

Yanbo Tang

Publications and source records attributed to Yanbo Tang.

12 recordsLinked to original sources

Efficient Stochastic Optimisation via Sequential Monte Carlo

The problem of optimising functions with intractable gradients frequently arises in machine learning and statistics, ranging from maximum marginal likelihood estimation procedures to fine-tuning of generative models. Stochastic approximation methods for this class of problems typically require inner sampling loops to obtain (biased) stochastic gradient estimates, which rapidly becomes computationally expensive. In this work, we develop sequential Monte Carlo (SMC) samplers for optimisation of functions with intractable gradients. Our approach replaces expensive inner sampling methods with efficient SMC approximations, which can result in significant computational gains. We establish convergence results for the basic recursions defined by our methodology which SMC samplers approximate. We demonstrate the effectiveness of our approach on the reward-tuning of energy-based models within various settings.

stat.ML

Post-reduction inference for confidence sets of models

Sparsity in a regression context makes the model itself an object of interest, pointing to a confidence set of models as the appropriate presentation of evidence. A difficulty in areas such as genomics, where the number of candidate variables is vast, arises from the need for preliminary reduction prior to the assessment of models. The present paper considers a resolution using inferential separations fundamental to the Fisherian approach to conditional inference, namely, the sufficiency/co-sufficiency separation, and the ancillary/co-ancillary separation. The advantage of these separations is that no direction for departure from any hypothesised model is needed, avoiding issues that would otherwise arise from using the same data for reduction and for model assessment. In idealised cases with no nuisance parameters, the separations extract all the information in the data solely for the purpose for which it is useful, without loss or redundancy. The extent to which estimation of nuisance parameters affects the idealised information extraction is illustrated in detail for the normal-theory linear regression model, extending immediately to a log-normal accelerated-life model for time-to-event outcomes. This idealised analysis provides insight into when sample-splitting is likely to perform as well as, or better than, the co-sufficient or ancillary tests, and when it may be unreliable. The considerations involved in extending the detailed implementation to canonical exponential-family and more general regression models are briefly discussed. As part of the analysis for the Gaussian model, we introduce a modified version of the refitted cross-validation estimator of Fan et al. (2012), whose distribution theory is tractable in the appropriate conditional sense.

math.ST

Monte Carlo and quasi-Monte Carlo integration for likelihood functions

We compare the integration error of Monte Carlo (MC) and quasi-Monte Carlo (QMC) methods for approximating the normalizing constant of posterior distributions and certain marginal likelihoods. In doing so, we characterize the dependency of the relative and absolute integration errors on the number of data points ($n$), the number of grid points ($m$) and the dimension of the integral ($p$). We find that if the dimension of the integral remains fixed as $n$ and $m$ tend to infinity, the scaling rate of the relative error of MC integration includes an additional $n^{1/2}\log(n)^{p/2}$ data-dependent factor, while for QMC this factor is $\log(n)^{p/2}$. In this scenario, QMC will outperform MC if $\log(m)^{p - 1/2}/\sqrt{mn\log(n)} < 1$, which differs from the usual result that QMC will outperform MC if $\log(m)^p/m^{1/2} < 1$.The accuracies of MC and QMC methods are also examined in the high-dimensional setting as $p \rightarrow \infty$, where MC gives more optimistic results as the scaling in dimension is slower than that of QMC when the Halton sequence is used to construct the low discrepancy grid; however both methods display poor dimensional scaling as expected. An additional contribution of this work is a bound on the high-dimensional scaling of the star discrepancy for the Halton sequence.

math.ST

Inference for overparametrized hierarchical Archimedean copulas

Hierarchical Archimedean copulas (HACs) are multivariate uniform distributions constructed by nesting Archimedean copulas into one another, and provide a flexible approach to modeling non-exchangeable data. However, this flexibility in the model structure may lead to over-fitting when the model estimation procedure is not performed properly. In this paper, we examine the problem of structure estimation and more generally on the selection of a parsimonious model from the hypothesis testing perspective. Formal tests for structural hypotheses concerning HACs have been lacking so far, most likely due to the restrictions on their associated parameter space which hinders the use of standard inference methodology. Building on previously developed asymptotic methods for these non-standard parameter spaces, we provide an asymptotic stochastic representation for the maximum likelihood estimators of (potentially) overparametrized HACs, which we then use to formulate a likelihood ratio test for certain common structural hypotheses. Additionally, we also derive analytical expressions for the first- and second-order partial derivatives of two-level HACs based on Clayton and Gumbel generators, as well as general numerical approximation schemes for the Fisher information matrix.

stat.ME

Variability of extragalactic X-ray jets on kiloparsec scales

Unexpectedly strong X-ray emission from extragalactic radio jets on kiloparsec scales has been one of the major discoveries of Chandra, the only X-ray observatory capable of sub-arcsecond-scale imaging. The origin of this X-ray emission, which appears as a second spectral component from that of the radio emission, has been debated for over two decades. The most commonly assumed mechanism is inverse Compton upscattering of the Cosmic Microwave Background (IC-CMB) by very low-energy electrons in a still highly relativistic jet. Under this mechanism, no variability in the X-ray emission is expected. Here we report the detection of X-ray variability in the large-scale jet population, using a novel statistical analysis of 53 jets with multiple Chandra observations. Taken as a population, we find that the distribution of p-values from a Poisson model is strongly inconsistent with steady emission, with a global p-value of 1.96e-4 under a Kolmogorov-Smirnov test against the expected Uniform (0,1) distribution. These results strongly imply that the dominant mechanism of X-ray production in kpc-scale jets is synchrotron emission by a second population of electrons reaching multi-TeV energies. X-ray variability on the time-scale of months to a few years implies extremely small emitting volumes much smaller than the cross-section of the jet.

astro-ph.HE

A Note on Monte Carlo Integration in High Dimensions

Monte Carlo integration is a commonly used technique to compute intractable integrals and is typically thought to perform poorly for very high-dimensional integrals. To show that this is not always the case, we examine Monte Carlo integration using techniques from the high-dimensional statistics literature by allowing the dimension of the integral to increase. In doing so, we derive non-asymptotic bounds for the relative and absolute error of the approximation for some general classes of functions through concentration inequalities. We provide concrete examples in which the magnitude of the number of points sampled needed to guarantee a consistent estimate vary between polynomial to exponential, and show that in theory arbitrarily fast or slow rates are possible. This demonstrates that the behaviour of Monte Carlo integration in high dimensions is not uniform. Through our methods we also obtain non-asymptotic confidence intervals for Monte Carlo estimate which are valid regardless of the number of points sampled.

stat.ME

Asymptotics of numerical integration for two-level mixed models

We study mixed models with a single grouping factor, where inference about unknown parameters requires optimizing a marginal likelihood defined by an intractable integral. Low-dimensional numerical integration techniques are regularly used to approximate these integrals, with inferences about parameters based on the resulting approximate marginal likelihood. For a generic class of mixed models that satisfy explicit regularity conditions, we derive the stochastic relative error rate incurred for both the likelihood and maximum likelihood estimator when adaptive numerical integration is used to approximate the marginal likelihood. We then specialize the analysis to well-specified generalized linear mixed models having exponential family response and multivariate Gaussian random effects, verifying that the regularity conditions hold, and hence that the convergence rates apply. We also prove that for models with likelihoods satisfying very weak concentration conditions that the maximum likelihood estimators from non-adaptive numerical integration approximations of the marginal likelihood are not consistent, further motivating adaptive numerical integration as the preferred tool for inference in mixed models. Code to reproduce the simulations in this paper is provided at https://github.com/awstringer1/aq-theory-paper-code.

stat.ME

Asymptotic Behaviour of the Modified Likelihood Root

We examine the normal approximation of the modified likelihood root, an inferential tool from higher-order asymptotic theory, for the linear exponential and location-scale family. We show that the $r^\star$ statistic can be thought of as a location and scale adjustment to the likelihood root $r$ up to $O_p(n^{-3/2})$, and more generally $r^\star$ can be expressed as a polynomial in $r$. We also show the linearity of the modified likelihood root in the likelihood root for these two families.

stat.ME

Laplace and Saddlepoint Approximations in High Dimensions

We examine the behaviour of the Laplace and saddlepoint approximations in the high-dimensional setting, where the dimension of the model is allowed to increase with the number of observations. Approximations to the joint density, the marginal posterior density and the conditional density are considered. Our results show that under the mildest assumptions on the model, the error of the joint density approximation is $O(p^4/n)$ if $p = o(n^{1/4})$ for the Laplace approximation and saddlepoint approximation, and $O(p^3/n)$ if $p = o(n^{1/3})$ under additional assumptions on the second derivative of the log-likelihood. Stronger results are obtained for the approximation to the marginal posterior density.

math.ST

Stochastic Convergence Rates and Applications of Adaptive Quadrature in Bayesian Inference

We provide the first stochastic convergence rates for a family of adaptive quadrature rules used to normalize the posterior distribution in Bayesian models. Our results apply to the uniform relative error in the approximate posterior density, the coverage probabilities of approximate credible sets, and approximate moments and quantiles, therefore guaranteeing fast asymptotic convergence of approximate summary statistics used in practice. The family of quadrature rules includes adaptive Gauss-Hermite quadrature, and we apply this rule in two challenging low-dimensional examples. Further, we demonstrate how adaptive quadrature can be used as a crucial component of a modern approximate Bayesian inference procedure for high-dimensional additive models. The method is implemented and made publicly available in the aghq package for the R language, available on CRAN.

stat.ME

General Behaviour of P-Values Under the Null and Alternative

Hypothesis testing results often rely on simple, yet important assumptions about the behaviour of the distribution of p-values under the null and the alternative. We examine tests for one dimensional parameters of interest that converge to a normal distribution, possibly in the presence of nuisance parameters, and characterize the distribution of the p-values using techniques from the higher order asymptotics literature. We show that commonly held beliefs regarding the distribution of p-values are misleading when the variance and location of the test statistic are not well-calibrated or when the higher order cumulants of the test statistic are not negligible. Corrected tests are proposed and are shown to perform better than their first order counterparts in certain settings.

math.ST