Searcharxiv⌕ Search

arXiv subjects

Robert Kohn

Publications and source records attributed to Robert Kohn.

At least 37 records · Page 2Linked to original sources

Flexible Variational Bayes based on a Copula of a Mixture

Variational Bayes methods approximate the posterior density by a family of tractable distributions whose parameters are estimated by optimisation. Variational approximation is useful when exact inference is intractable or very costly. Our article develops a flexible variational approximation based on a copula of a mixture, which is implemented by combining boosting, natural gradient, and a variance reduction method. The efficacy of the approach is illustrated by using simulated and real datasets to approximate multimodal, skewed and heavy-tailed posterior distributions, including an application to Bayesian deep feedforward neural network regression models.

stat.CO↗

Adaptively switching between a particle marginal Metropolis-Hastings and a particle Gibbs kernel in SMC$^2$

Sequential Monte Carlo squared (SMC$^2$; Chopin et al., 2012) methods can be used to sample from the exact posterior distribution of intractable likelihood state space models. These methods are the SMC analogue to particle Markov chain Monte Carlo (MCMC; Andrieu et al., 2010) and rely on particle MCMC kernels to mutate the particles at each iteration. Two options for the particle MCMC kernels are particle marginal Metropolis-Hastings (PMMH) and particle Gibbs (PG). We introduce a method to adaptively select the particle MCMC kernel at each iteration of SMC$^2$, with a particular focus on switching between a PMMH and PG kernel. The resulting method can significantly improve the efficiency of SMC$^2$ compared to using a fixed particle MCMC kernel throughout the algorithm. Code for our methods is available at https://github.com/imkebotha/kernel_switching_smc2.

stat.CO↗

The Correlated Particle Hybrid Sampler for State Space Models

Particle Markov Chain Monte Carlo (PMCMC) is a general computational approach to Bayesian inference for general state space models. Our article scales up PMCMC in terms of the number of observations and parameters by generating the parameters that are highly correlated with the states \lq integrated out\rq{} in a pseudo marginal step; the rest of the parameters are generated conditional on the states. The novel contribution of our article is to make the pseudo-marginal step much more efficient by positively correlating the numerator and denominator in the Metropolis-Hastings acceptance probability. This is done in a novel way by expressing the target density of the PMCMC in terms of the basic uniform or normal random numbers used in the sequential Monte Carlo algorithm instead of the standard way in terms of state particles. We also show that the new sampler combines and generalizes two separate particle MCMC approaches: particle Gibbs and the correlated pseudo marginal Metropolis-Hastings. We investigate the performance of the hybrid sampler empirically by applying it to univariate and multivariate stochastic volatility models having both a large number of parameters and a large number of latent states and show that it is much more efficient than competing PMCMC methods.

stat.ME↗

Bayesian Inference for Evidence Accumulation Models with Regressors

Evidence accumulation models (EAMs) are an important class of cognitive models used to analyze both response time and response choice data recorded from decision-making tasks. Developments in estimation procedures have helped EAMs become important both in basic scientific applications and solution-focussed applied work. Hierarchical Bayesian estimation frameworks for the linear ballistic accumulator model (LBA) and the diffusion decision model (DDM) have been widely used, but still suffer from some key limitations, particularly for large sample sizes, for models with many parameters, and when linking decision-relevant covariates to model parameters. We extend upon previous work with methods for estimating the LBA and DDM in hierarchical Bayesian frameworks that include random effects which are correlated between people, and include regression-model links between decision-relevant covariates and model parameters. Our methods work equally well in cases where the covariates are measured once per person (e.g., personality traits or psychological tests) or once per decision (e.g., neural or physiological data). We provide methods for exact Bayesian inference, using particle-based MCMC, and also approximate methods based on variational Bayesian (VB) inference. The VB methods are sufficiently fast and efficient that they can address large-scale estimation problems, such as with very large data sets. We evaluate the performance of these methods in applications to data from three existing experiments. Detailed algorithmic implementations and code are freely available for all methods.

stat.ME↗

Particle Mean Field Variational Bayes

The Mean Field Variational Bayes (MFVB) method is one of the most computationally efficient techniques for Bayesian inference. However, its use has been restricted to models with conjugate priors or those that require analytical calculations. This paper proposes a novel particle-based MFVB approach that greatly expands the applicability of the MFVB method. We establish the theoretical basis of the new method by leveraging the connection between Wasserstein gradient flows and Langevin diffusion dynamics, and demonstrate the effectiveness of this approach using Bayesian logistic regression, stochastic volatility, and deep neural networks.

stat.CO↗

The Block-Correlated Pseudo Marginal Sampler for State Space Models

Particle Marginal Metropolis-Hastings (PMMH) is a general approach to Bayesian inference when the likelihood is intractable, but can be estimated unbiasedly. Our article develops an efficient PMMH method that scales up better to higher dimensional state vectors than previous approaches. The improvement is achieved by the following innovations. First, the trimmed mean of the unbiased likelihood estimates of the multiple particle filters is used. Second, a novel block version of PMMH that works with multiple particle filters is proposed. Third, the article develops an efficient auxiliary disturbance particle filter, which is necessary when the bootstrap disturbance filter is inefficient, but the state transition density cannot be expressed in closed form. Fourth, a novel sorting algorithm, which is as effective as previous approaches but significantly faster than them, is developed to preserve the correlation between the logs of the likelihood estimates at the current and proposed parameter values. The performance of the sampler is investigated empirically by applying it to non-linear Dynamic Stochastic General Equilibrium models with relatively high state dimensions and with intractable state transition densities and to multivariate stochastic volatility in the mean models. Although our focus is on applying the method to state space models, the approach will be useful in a wide range of applications such as large panel data models and stochastic differential equation models with mixed effects.

stat.CO↗

Reliable Bayesian Inference in Misspecified Models

We provide a general solution to a fundamental open problem in Bayesian inference, namely poor uncertainty quantification, from a frequency standpoint, of Bayesian methods in misspecified models. While existing solutions are based on explicit Gaussian approximations of the posterior, or computationally onerous post-processing procedures, we demonstrate that correct uncertainty quantification can be achieved by replacing the usual posterior with an intuitive approximate posterior. Critically, our solution is applicable to likelihood-based, and generalized, posteriors as well as cases where the likelihood is intractable and must be estimated. We formally demonstrate the reliable uncertainty quantification of our proposed approach, and show that valid uncertainty quantification is not an asymptotic result but occurs even in small samples. We illustrate this approach through a range of examples, including linear, and generalized, mixed effects models.

stat.ME↗

Structured variational approximations with skew normal decomposable graphical models

Although there is much recent work developing flexible variational methods for Bayesian computation, Gaussian approximations with structured covariance matrices are often preferred computationally in high-dimensional settings. This paper considers approximate inference methods for complex latent variable models where the posterior is close to Gaussian, but with some skewness in the posterior marginals. We consider skew decomposable graphical models (SDGMs), which are based on the closed skew normal family of distributions, as variational approximations. These approximations can reflect the true posterior conditional independence structure and capture posterior skewness. Different parametrizations are explored for this variational family, and the speed of convergence and quality of the approximation can depend on the parametrization used. To increase flexibility, implicit copula SDGM approximations are also developed, where elementwise transformations of an approximately standardized SDGM random vector are considered. Our parametrization of the implicit copula approximation is novel, even in the special case of a Gaussian approximation. Performance of the methods is examined in a number of real examples involving generalized linear mixed models and state space models, and we conclude that our copula approaches are most accurate, but that the SDGM methods are often nearly as good and have lower computational demands.

stat.CO↗

On approximating copulas by finite mixtures

Copulas are now frequently used to construct or estimate multivariate distributions because of their ability to take into account the multivariate dependence of the different variables while separately specifying marginal distributions. Copula based multivariate models can often also be more parsimonious than fitting a flexible multivariate model, such as a mixture of normals model, directly to the data. However, to be effective, it is imperative that the family of copula models considered is sufficiently flexible. Although finite mixtures of copulas have been used to construct flexible families of copulas, their approximation properties are not well understood and we show that natural candidates such as mixtures of elliptical copulas and mixtures of Archimedean copulas cannot approximate a general copula arbitrarily well. Our article develops fundamental tools for approximating a general copula arbitrarily well by a copulas based on finite mixtures. We show the asymptotic properties as well as illustrate the advantages of our methodology empirically on a financial data set and on some artificial data.

stat.ME↗

Automatically adapting the number of state particles in SMC$^2$

Sequential Monte Carlo squared (SMC$^2$) methods can be used for parameter inference of intractable likelihood state-space models. These methods replace the likelihood with an unbiased particle filter estimator, similarly to particle Markov chain Monte Carlo (MCMC). As with particle MCMC, the efficiency of SMC$^2$ greatly depends on the variance of the likelihood estimator, and therefore on the number of state particles used within the particle filter. We introduce novel methods to adaptively select the number of state particles within SMC$^2$ using the expected squared jumping distance to trigger the adaptation, and modifying the exchange importance sampling method of \citet{Chopin2012a} to replace the current set of state particles with the new set of state particles. The resulting algorithm is fully automatic, and can significantly improve current methods. Code for our methods is available at https://github.com/imkebotha/adaptive-exact-approximate-smc.

stat.CO↗

Dynamic Mixture of Experts Models for Online Prediction

A mixture of experts models the conditional density of a response variable using a mixture of regression models with covariate-dependent mixture weights. We extend the finite mixture of experts model by allowing the parameters in both the mixture components and the weights to evolve in time by following random walk processes. Inference for time-varying parameters in richly parameterized mixture of experts models is challenging. We propose a sequential Monte Carlo algorithm for online inference and based on a tailored proposal distribution built on ideas from linear Bayes methods and the EM algorithm. The method gives a unified treatment for mixtures with time-varying parameters, including the special case of static parameters. We assess the properties of the method on simulated data and on industrial data where the aim is to predict software faults in a continuously upgraded large-scale software project.

stat.CO↗

Spectral Subsampling MCMC for Stationary Multivariate Time Series with Applications to Vector ARTFIMA Processes

Spectral subsampling MCMC was recently proposed to speed up Markov chain Monte Carlo (MCMC) for long stationary univariate time series by subsampling periodogram observations in the frequency domain. This article extends the approach to multivariate time series using a multivariate generalisation of the Whittle likelihood. To assess the computational gains from spectral subsampling in challenging problems, a multivariate generalisation of the autoregressive tempered fractionally integrated moving average model (ARTFIMA) is introduced and some of its properties derived. Bayesian inference based on the Whittle likelihood is demonstrated to be a fast and accurate alternative to the exact time domain likelihood. Spectral subsampling is shown to provide up to two orders of magnitude additional speed-up, while retaining MCMC sampling efficiency and accuracy, compared to spectral methods using the full dataset. Keywords: Bayesian, Markov chain Monte Carlo, Semi-long memory, Spectral analysis, Whittle likelihood.

stat.ME↗

An energy minimization approach to twinning with variable volume fraction

In materials that undergo martensitic phase transformation, macroscopic loading often leads to the creation and/or rearrangement of elastic domains. This paper considers an example {involving} a single-crystal slab made from two martensite variants. When the slab is made to bend, the two variants form a characteristic microstructure that we like to call ``twinning with variable volume fraction.'' Two 1996 papers by Chopra et. al. explored this example using bars made from InTl, providing considerable detail about the microstructures they observed. Here we offer an energy-minimization-based model that is motivated by their account. It uses geometrically linear elasticity, and treats the phase boundaries as sharp interfaces. For simplicity, rather than model the experimental forces and boundary conditions exactly, we consider certain Dirichlet or Neumann boundary conditions whose effect is to require bending. This leads to certain nonlinear (and nonconvex) variational problems that represent the minimization of elastic plus surface energy (and the work done by the load, in the case of a Neumann boundary condition). Our results identify how the minimum value of each variational problem scales with respect to the surface energy density. The results are established by proving upper and lower bounds that scale the same way. The upper bounds are ansatz-based, providing full details about some (nearly) optimal microstructures. The lower bounds are ansatz-free, so they explain why no other arrangement of the two phases could be significantly better.

math.AP↗

Robust Particle Density Tempering for State Space Models

Density tempering (also called density annealing) is a sequential Monte Carlo approach to Bayesian inference for general state models; it is an alternative to Markov chain Monte Carlo. When applied to state space models, it moves a collection of parameters and latent states (which are called particles) through a number of stages, with each stage having its own target distribution. The particles are initially generated from a distribution that is easy to sample from, e.g. the prior; the target at the final stage is the posterior distribution. Tempering is usually carried out either in batch mode, involving all the data at each stage, or sequentially with observations added at each stage, which is called data tempering. Our paper proposes efficient Markov moves for generating the parameters and states for each stage of particle based density tempering. This allows the proposed SMC methods to increase (scale up) the number of parameters and states that can be handled. Most of the current literature uses a pseudo-marginal Markov move step with the states integrated out, and the parameters generated by a random walk proposal; although this strategy is general, it is very inefficient when the states or parameters are high dimensional. We also build on the work of Dufays (2016) and make data tempering more robust to outliers and structural changes for models with intractable likelihoods by adding batch tempering at each stage. The performance of the proposed methods is evaluated using univariate stochastic volatility models with outliers and structural breaks and high dimensional factor stochastic volatility models having both many parameters and many latent states.

stat.ME↗

Time-evolving psychological processes over repeated decisions

Many psychological experiments have subjects repeat a task to gain the statistical precision required to test quantitative theories of psychological performance. In such experiments, time-on-task can have sizable effects on performance, changing the psychological processes under investigation. Most research has either ignored these changes, treating the underlying process as static, or sacrificed some psychological content of the models for statistical simplicity. We use particle Markov chain Monte-Carlo methods to study psychologically plausible time-varying changes in model parameters. Using data from three highly-cited experiments we find strong evidence in favor of a hidden Markov switching process as an explanation of time-varying effects. This embodies the psychological assumption of "regime switching", with subjects alternating between different cognitive states representing different modes of decision-making. The switching model explains key long- and short-term dynamic effects in the data. The central idea of our approach can be applied quite generally to quantitative psychological theories, beyond the models and data sets that we investigate.

stat.AP↗

Efficient Selection Between Hierarchical Cognitive Models: Cross-validation With Variational Bayes

Model comparison is the cornerstone of theoretical progress in psychological research. Common practice overwhelmingly relies on tools that evaluate competing models by balancing in-sample descriptive adequacy against model flexibility, with modern approaches advocating the use of marginal likelihood for hierarchical cognitive models. Cross-validation is another popular approach but its implementation has remained out of reach for cognitive models evaluated in a Bayesian hierarchical framework, with the major hurdle being prohibitive computational cost. To address this issue, we develop novel algorithms that make variational Bayes (VB) inference for hierarchical models feasible and computationally efficient for complex cognitive models of substantive theoretical interest. It is well known that VB produces good estimates of the first moments of the parameters which gives good predictive densities estimates. We thus develop a novel VB algorithm with Bayesian prediction as a tool to perform model comparison by cross-validation, which we refer to as CVVB. In particular, the CVVB can be used as a model screening device that quickly identifies bad models. We demonstrate the utility of CVVB by revisiting a classic question in decision making research: what latent components of processing drive the ubiquitous speed-accuracy tradeoff? We demonstrate that CVVB strongly agrees with model comparison via marginal likelihood yet achieves the outcome in much less time. Our approach brings cross-validation within reach of theoretically important psychological models, and makes it feasible to compare much larger families of hierarchically specified cognitive models than has previously been possible.

stat.AP↗

Modelling age-related changes in executive functions of soccer players

The widespread popularity of soccer across the globe has turned it into a multi-billion dollar industry. As a result, most professional clubs actively engage in talent identification and development programmes. Contemporary research has generally supported the use of executive functions - a class of neuropsychological processes responsible for cognitive behaviours - in predicting a soccer player's future success. However, studies on the developmental evolution of executive functions have yielded differing results in their structural form (such as inverted U-shapes, or otherwise). This article presents the first analysis of changes in the domain-generic and domain-specific executive functions based on longitudinal data measured on elite German soccer players. Results obtained from a latent variable model show that these executive functions experience noticeable growth from late childhood until pre-adolescence, but remain fairly stable in later growth stages. As a consequence, our results suggest that the role of executive functions in facilitating talent identification may have been overly emphasised.

stat.AP↗

Variational Approximation of Factor Stochastic Volatility Models

Estimation and prediction in high dimensional multivariate factor stochastic volatility models is an important and active research area because such models allow a parsimonious representation of multivariate stochastic volatility. Bayesian inference for factor stochastic volatility models is usually done by Markov chain Monte Carlo methods, often by particle Markov chain Monte Carlo, which are usually slow for high dimensional or long time series because of the large number of parameters and latent states involved. Our article makes two contributions. The first is to propose fast and accurate variational Bayes methods to approximate the posterior distribution of the states and parameters in factor stochastic volatility models. The second contribution is to extend this batch methodology to develop fast sequential variational updates for prediction as new observations arrive. The methods are applied to simulated and real datasets and shown to produce good approximate inference and prediction compared to the latest particle Markov chain Monte Carlo approaches, but are much faster.

stat.CO↗