SearcharxivSearch

arXiv subjects

Kasper Kristensen

Publications and source records attributed to Kasper Kristensen.

8 recordsLinked to original sources

Leveraging Sparsity to Improve No-U-Turn Sampling Efficiency for Hierarchical Bayesian Models

Analysts routinely use Bayesian hierarchical models to understand natural processes. The no-U-turn sampler (NUTS) is the most widely used algorithm to sample high-dimensional, continuously differentiable models. But NUTS is slowed by high correlations, especially in high dimensions, limiting the complexity of applied analyses. Here we introduce Sparse NUTS (SNUTS), which preconditions (decorrelates and descales) posteriors using a sparse precision matrix ($Q$). We use Template Model Builder (TMB) to efficiently compute $Q$ from the mode of the Laplace approximation to the marginal posterior, then pass the preconditioned posterior to NUTS through the Bayesian software Stan for sampling. We apply SNUTS to seventeen diverse case studies to demonstrate that preconditioning with $Q$ converges one to two orders of magnitude faster than Stan's industry standard diagonal or dense preconditioners. SNUTS also outperforms preconditioning with the inverse of the covariance estimated with Pathfinder variational inference. SNUTS does not improve sampling efficiency for models with the highly varying curvature found in funnels, wide tails, or multiple modes. SNUTS is most advantageous, and can be scaled beyond $10^4$ parameters, in the presence of high dimensionality, sparseness, and high correlations, all of which are widespread in applied statistics. An open-source implementation of SNUTS is provided in the R package SparseNUTS.

stat.CO

Inference in stochastic differential equations using the Laplace approximation: Demonstration and examples

Stochastic differential equations are a natural framework for dynamic systems and time series in ecology, because they allow for non-linear first-principle knowledge and uncertainty in the dynamics, and can be combined with measurement errors. However, estimation methods are often technically and computationally challenging. Here, we demonstrate that the Laplace approximation is useful for estimating states and parameters in these models, when done correctly. We give special attention to non-linear dynamics, state-dependent noise intensities, and non-Gaussian measurement errors. Our technique adds states between times of observations, approximates transition densities using discretization methods - in the simplest case, the Euler-Maruyama method - and eliminates unobserved states using the Laplace approximation. We demonstrate that consistency requires a particular form of the approximation, and provide different approaches to implementation. Using simulated case studies, we demonstrate that transition probabilities are well approximated, that inference is computationally feasible, and that the framework leads to simple and flexible implementations.

stat.ME

Computing Sparse Jacobians and Hessians Using Algorithmic Differentiation

Stochastic scientific models and machine learning optimization estimators have a large number of variables; hence computing large sparse Jacobians and Hessians is important. Algorithmic differentiation (AD) greatly reduces the programming effort required to obtain the sparsity patterns and values for these matrices. We present forward, reverse, and subgraph methods for computing sparse Jacobians and Hessians. Special attention is given the the subgraph method because it is new. The coloring and compression steps are not necessary when computing sparse Jacobians and Hessians using subgraphs. Complexity analysis shows that for some problems the subgraph method is expected to be much faster. We compare C++ operator overloading implementations of the methods in the ADOL-C and CppAD software packages using some of the MINPACK-2 test problems. The experiments are set up in a way that makes them easy to run on different hardware, different systems, different compilers, other test problem and other AD packages. The setup time is the time to record the graph, compute sparsity, coloring, compression, and optimization of the graph. If the setup is necessary for each evaluation, the subgraph implementation has similar run times for sparse Jacobians and faster run times for sparse Hessians.

cs.MS

The Multiplicative Mixed Model with the mumm R package as a General and Easy Random Interaction Model Tool

Multiplicative mixed models can be applied in a wide range of scientific disciplines, since they are relevant in every situation where an interaction between a fixed effect and a random effect is present. Until now, no R package has been published, which can fit this type of models. The lack of user-friendly open source tools to fit these models, is the main reason that the models are not used as often as they could or should be. In this paper we introduce the user-friendly R package $\mathsf{mumm}$ for fitting multiplicative mixed models in a time-efficient manner. To illustrate the interpretation of the multiplicative term, we provide four data analysis examples, where the model is fitted to data sets that stem from studies in sensometrics, agriculture and medicine. With these examples it is shown that the statistical inference can be improved by using a multiplicative mixed model, instead of a linear mixed model which is usually employed.

stat.CO

Convergence of coupled cluster perturbation theory

The convergence of a recently proposed coupled cluster (CC) family of perturbation series [Eriksen, J. J. et al., J. Chem. Phys. 140, 064108 (2014)], in which the energetic difference between two CC models - a low-level parent and a high-level target model - is expanded in orders of the Møller-Plesset (MP) fluctuation potential, is investigated for four prototypical closed-shell systems (Ne, singlet methylene, distorted HF, and the fluoride anion) in standard and augmented basis sets. In these investigations, energy corrections of the various series have been calculated to high orders and their convergence radii determined by probing for possible front- and back-door intruder states, the existence of which would make the series divergent. In summary, we conclude how it is primarily the choice of target state, and not the choice of parent state, which ultimately governs the convergence behavior of a given series. For example, restricting the target state to, say, triple or quadruple excitations might remove intruders present in series that target the full configuration interaction (FCI) limit, such as the standard MP series. Furthermore, we find that whereas a CC perturbation series might converge within standard correlation consistent basis sets, it may start to diverge whenever these become augmented by diffuse functions, similar to the MP case. However, unlike for the MP case, such potential divergences are not found to invalidate the practical use of the low-order corrections of the CC perturbation series.

physics.chem-ph

CC2 oscillator strengths within the local framework for calculating excitation energies (LoFEx)

In a recent work [Baudin and Kristensen, J. Chem. Phys. 144, 224106 (2016)], we introduced a local framework for calculating excitation energies (LoFEx), based on second-order approximated coupled cluster (CC2) linear-response theory. LoFEx is a black-box method in which a reduced excitation orbital space (XOS) is optimized to provide coupled cluster (CC) excitation energies at a reduced computational cost. In this article, we present an extension of the LoFEx algorithm to the calculation of CC2 oscillator strengths. Two different strategies are suggested, in which the size of the XOS is determined based on the excitation energy or the oscillator strength of the targeted transitions. The two strategies are applied to a set of medium-sized organic molecules in order to assess both the accuracy and the computational cost of the methods. The results show that CC2 excitation energies and oscillator strengths can be calculated at a reduced computational cost, provided that the targeted transitions are local compared to the size of the molecule. To illustrate the potential of LoFEx for large molecules, both strategies have been successfully applied to the lowest transition of the bivalirudin molecule (4255 basis functions) and compared with time-dependent density functional theory.

physics.chem-ph

A view on coupled cluster perturbation theory using a bivariational Lagrangian formulation

We consider two distinct coupled cluster (CC) perturbation series that both expand the difference between the energies of the CCSD (CC with single and double excitations) and CCSDT (CC with single, double, and triple excitations) models in orders of the Møller-Plesset fluctuation potential. We initially introduce the E-CCSD(T-$n$) series, in which the CCSD amplitude equations are satisfied at the expansion point, and compare it to the recently developed CCSD(T-$n$) series [J. Chem. Phys. 140, 064108 (2014)], in which not only the CCSD amplitude, but also the CCSD multiplier equations are satisfied at the expansion point. The computational scaling is similar for the two series, and both are term-wise size extensive with a formal convergence towards the CCSDT target energy. However, the two series are different, and the CCSD(T-$n$) series is found to exhibit a more rapid convergence up through the series, which we trace back to the fact that more information at the expansion point is utilized than for the E-CCSD(T-$n$) series. The present analysis can be generalized to any perturbation expansion representing the difference between a parent CC model and a higher-level target CC model. In general, we demonstrate that, whenever the parent parameters depend upon the perturbation operator, a perturbation expansion of the CC energy (where only parent amplitudes are used) differs from a perturbation expansion of the CC Lagrangian (where both parent amplitudes and parent multipliers are used). For the latter case, the bivariational Lagrangian formulation becomes more than a convenient mathematical tool, since it facilitates a different and faster convergent perturbation series than the simpler energy-based expansion.

physics.chem-ph

TMB: Automatic Differentiation and Laplace Approximation

TMB is an open source R package that enables quick implementation of complex nonlinear random effect (latent variable) models in a manner similar to the established AD Model Builder package (ADMB, admb-project.org). In addition, it offers easy access to parallel computations. The user defines the joint likelihood for the data and the random effects as a C++ template function, while all the other operations are done in R; e.g., reading in the data. The package evaluates and maximizes the Laplace approximation of the marginal likelihood where the random effects are automatically integrated out. This approximation, and its derivatives, are obtained using automatic differentiation (up to order three) of the joint likelihood. The computations are designed to be fast for problems with many random effects (~10^6) and parameters (~10^3). Computation times using ADMB and TMB are compared on a suite of examples ranging from simple models to large spatial models where the random effects are a Gaussian random field. Speedups ranging from 1.5 to about 100 are obtained with increasing gains for large problems. The package and examples are available at http://tmb-project.org.

stat.CO