Searcharxiv⌕ Search

arXiv subjects

David Eriksson

Publications and source records attributed to David Eriksson.

34 records · Page 2Linked to original sources

High-Dimensional Bayesian Optimization with Sparse Axis-Aligned Subspaces

Bayesian optimization (BO) is a powerful paradigm for efficient optimization of black-box objective functions. High-dimensional BO presents a particular challenge, in part because the curse of dimensionality makes it difficult to define -- as well as do inference over -- a suitable class of surrogate models. We argue that Gaussian process surrogate models defined on sparse axis-aligned subspaces offer an attractive compromise between flexibility and parsimony. We demonstrate that our approach, which relies on Hamiltonian Monte Carlo for inference, can rapidly identify sparse subspaces relevant to modeling the unknown objective function, enabling sample-efficient high-dimensional BO. In an extensive suite of experiments comparing to existing methods for high-dimensional BO we demonstrate that our algorithm, Sparse Axis-Aligned Subspace BO (SAASBO), achieves excellent performance on several synthetic and real-world problems without the need to set problem-specific hyperparameters.

cs.LG↗

A Nonmyopic Approach to Cost-Constrained Bayesian Optimization

Bayesian optimization (BO) is a popular method for optimizing expensive-to-evaluate black-box functions. BO budgets are typically given in iterations, which implicitly assumes each evaluation has the same cost. In fact, in many BO applications, evaluation costs vary significantly in different regions of the search space. In hyperparameter optimization, the time spent on neural network training increases with layer size; in clinical trials, the monetary cost of drug compounds vary; and in optimal control, control actions have differing complexities. Cost-constrained BO measures convergence with alternative cost metrics such as time, money, or energy, for which the sample efficiency of standard BO methods is ill-suited. For cost-constrained BO, cost efficiency is far more important than sample efficiency. In this paper, we formulate cost-constrained BO as a constrained Markov decision process (CMDP), and develop an efficient rollout approximation to the optimal CMDP policy that takes both the cost and future iterations into account. We validate our method on a collection of hyperparameter optimization problems as well as a sensor set selection application.

cs.LG↗

Scalable Constrained Bayesian Optimization

The global optimization of a high-dimensional black-box function under black-box constraints is a pervasive task in machine learning, control, and engineering. These problems are challenging since the feasible set is typically non-convex and hard to find, in addition to the curses of dimensionality and the heterogeneity of the underlying functions. In particular, these characteristics dramatically impact the performance of Bayesian optimization methods, that otherwise have become the de facto standard for sample-efficient optimization in unconstrained settings, leaving practitioners with evolutionary strategies or heuristics. We propose the scalable constrained Bayesian optimization (SCBO) algorithm that overcomes the above challenges and pushes the applicability of Bayesian optimization far beyond the state-of-the-art. A comprehensive experimental evaluation demonstrates that SCBO achieves excellent results on a variety of benchmarks. To this end, we propose two new control problems that we expect to be of independent value for the scientific community.

cs.LG↗

Fast Matrix Square Roots with Applications to Gaussian Processes and Bayesian Optimization

Matrix square roots and their inverses arise frequently in machine learning, e.g., when sampling from high-dimensional Gaussians $\mathcal{N}(\mathbf 0, \mathbf K)$ or whitening a vector $\mathbf b$ against covariance matrix $\mathbf K$. While existing methods typically require $O(N^3)$ computation, we introduce a highly-efficient quadratic-time algorithm for computing $\mathbf K^{1/2} \mathbf b$, $\mathbf K^{-1/2} \mathbf b$, and their derivatives through matrix-vector multiplication (MVMs). Our method combines Krylov subspace methods with a rational approximation and typically achieves $4$ decimal places of accuracy with fewer than $100$ MVMs. Moreover, the backward pass requires little additional computation. We demonstrate our method's applicability on matrices as large as $50,\!000 \times 50,\!000$ - well beyond traditional methods - with little approximation error. Applying this increased scalability to variational Gaussian processes, Bayesian optimization, and Gibbs sampling results in more powerful models with higher accuracy.

cs.LG↗

Efficient Rollout Strategies for Bayesian Optimization

Bayesian optimization (BO) is a class of sample-efficient global optimization methods, where a probabilistic model conditioned on previous observations is used to determine future evaluations via the optimization of an acquisition function. Most acquisition functions are myopic, meaning that they only consider the impact of the next function evaluation. Non-myopic acquisition functions consider the impact of the next $h$ function evaluations and are typically computed through rollout, in which $h$ steps of BO are simulated. These rollout acquisition functions are defined as $h$-dimensional integrals, and are expensive to compute and optimize. We show that a combination of quasi-Monte Carlo, common random numbers, and control variates significantly reduce the computational burden of rollout. We then formulate a policy-search based approach that removes the need to optimize the rollout acquisition function. Finally, we discuss the qualitative behavior of rollout policies in the setting of multi-modal objectives and model error.

cs.LG↗

Scalable Global Optimization via Local Bayesian Optimization

Bayesian optimization has recently emerged as a popular method for the sample-efficient optimization of expensive black-box functions. However, the application to high-dimensional problems with several thousand observations remains challenging, and on difficult problems Bayesian optimization is often not competitive with other paradigms. In this paper we take the view that this is due to the implicit homogeneity of the global probabilistic models and an overemphasized exploration that results from global acquisition. This motivates the design of a local probabilistic approach for global optimization of large-scale high-dimensional problems. We propose the $\texttt{TuRBO}$ algorithm that fits a collection of local models and performs a principled global allocation of samples across these models via an implicit bandit approach. A comprehensive evaluation demonstrates that $\texttt{TuRBO}$ outperforms state-of-the-art methods from machine learning and operations research on problems spanning reinforcement learning, robotics, and the natural sciences.

cs.LG↗

pySOT and POAP: An event-driven asynchronous framework for surrogate optimization

This paper describes Plumbing for Optimization with Asynchronous Parallelism (POAP) and the Python Surrogate Optimization Toolbox (pySOT). POAP is an event-driven framework for building and combining asynchronous optimization strategies, designed for global optimization of expensive functions where concurrent function evaluations are useful. POAP consists of three components: a worker pool capable of function evaluations, strategies to propose evaluations or other actions, and a controller that mediates the interaction between the workers and strategies. pySOT is a collection of synchronous and asynchronous surrogate optimization strategies, implemented in the POAP framework. We support the stochastic RBF method by Regis and Shoemaker along with various extensions of this method, and a general surrogate optimization strategy that covers most Bayesian optimization methods. We have implemented many different surrogate models, experimental designs, acquisition functions, and a large set of test problems. We make an extensive comparison between synchronous and asynchronous parallelism and find that the advantage of asynchronous computation increases as the variance of the evaluation time or number of processors increases. We observe a close to linear speed-up with 4, 8, and 16 processors in both the synchronous and asynchronous setting.

math.OC↗

Scaling Gaussian Process Regression with Derivatives

Gaussian processes (GPs) with derivatives are useful in many applications, including Bayesian optimization, implicit surface reconstruction, and terrain reconstruction. Fitting a GP to function values and derivatives at $n$ points in $d$ dimensions requires linear solves and log determinants with an ${n(d+1) \times n(d+1)}$ positive definite matrix -- leading to prohibitive $\mathcal{O}(n^3d^3)$ computations for standard direct methods. We propose iterative solvers using fast $\mathcal{O}(nd)$ matrix-vector multiplications (MVMs), together with pivoted Cholesky preconditioning that cuts the iterations to convergence by several orders of magnitude, allowing for fast kernel learning and prediction. Our approaches, together with dimensionality reduction, enables Bayesian optimization with derivatives to scale to high-dimensional problems and large evaluation budgets.

cs.LG↗

Scalable Log Determinants for Gaussian Process Kernel Learning

For applications as varied as Bayesian neural networks, determinantal point processes, elliptical graphical models, and kernel learning for Gaussian processes (GPs), one must compute a log determinant of an $n \times n$ positive definite matrix, and its derivatives - leading to prohibitive $\mathcal{O}(n^3)$ computations. We propose novel $\mathcal{O}(n)$ approaches to estimating these quantities from only fast matrix vector multiplications (MVMs). These stochastic approximations are based on Chebyshev, Lanczos, and surrogate models, and converge quickly even for kernel matrices that have challenging spectra. We leverage these approximations to develop a scalable Gaussian process approach to kernel learning. We find that Lanczos is generally superior to Chebyshev for kernel learning, and that a surrogate approach can be highly efficient and accurate with popular kernels.

stat.ML↗

2HDMC - Two-Higgs-Doublet Model Calculator

This manual describes the public code 2HDMC which can be used to perform calculations in a general, CP-conserving, two-Higgs-doublet model (2HDM). The program features simple conversion between different parametrizations of the 2HDM potential, a flexible Yukawa sector specification with choices of different Z_2-symmetries or more general couplings, a tree-level decay library including all two-body - and some three-body - decay modes for the Higgs bosons, and the possibility to calculate observables of interest for constraining the 2HDM parameter space, as well as theoretical constraints from positivity and unitarity. The latest version of the 2HDMC code and full documentation is available from: http://www.isv.uu.se/thep/MC/2HDMC

hep-ph↗

PYBBWH: A program for associated charged Higgs and W boson production

The Monte Carlo program, PYBBWH, is an implementation of the associated production of a charged Higgs and a W boson from bb fusion in a general Two-Higgs-Doublet model for both CP-conserving and CP-violating couplings. It is implemented as a external process to Pythia 6. The code can be downloaded from http://www.isv.uu.se/thep/MC/pybbwh/

hep-ph↗

Colour rearrangements in B-meson decays

We present a new model, based on colour rearrangements, which at the same time can describe both hidden and open charm production in B-meson decays. The model is successfully compared to both inclusive decays, such as B to J/psi X and B to D_s X, as well as exclusive ones, such as B to J/psi K^(*) and B to D^(*) D^(*)K. It also gives a good description of the momentum distribution of direct J/psi's, especially in the low-momentum region, which earlier has been claimed as a possible signal for new exotic states.

hep-ph↗

Associated charged Higgs and W boson production in the MSSM at the CERN Large Hadron Collider

We investigate the viability of observing charged Higgs bosons (H^+/-) produced in association with W bosons at the CERN Large Hadron Collider, using the leptonic decay H^+ -> tau^+ nu_tau and hadronic W-decay, within different scenarios of the Minimal Supersymmetric Standard Model (MSSM) with both real and complex parameters. Performing a parton level study we show how the irreducible Standard Model background from W+2 jets can be controlled by applying appropriate cuts and find that the size of a possible signal depends on the cuts needed to suppress QCD backgrounds and misidentifications. In the standard maximal mixing scenario of the MSSM we find a viable signal for large tan(beta) and intermediate H^+/- masses (~m_t) when using optimistic cuts whereas for more pessimistic ones we only find a viable signal for very large tan(beta) (>~50). We have also investigated a special class of MSSM scenarios with large mass-splittings among the heavy Higgs bosons where the cross-section can be resonantly enhanced by factors up to one hundred, with a strong dependence on the CP-violating phases. Even so we find that the signal after cuts remains small except for small masses (~< m_t) with optimistic cuts. Finally, in all the scenarios we have investigated we have only found small CP-asymmetries.

hep-ph↗

New angles on top quark decay to a charged Higgs

To properly discover a charged Higgs Boson ($H^\pm$) requires its spin and couplings to be determined. We investigate how to utilize $\ttbar$ spin correlations to analyze the $H^\pm$ couplings in the decay $t\to bH^+\to bτ^+ν_τ$. Within the framework of a general Two-Higgs-Doublet Model, we obtain results on the spin analyzing coefficients for this decay and study in detail its spin phenomenology, focusing on the limits of large and small values for $\tanβ$. Using a Monte Carlo approach to simulate full hadron-level events, we evaluate systematically how the $H^\pm\toτ^\pmν_τ$ decay mode can be used for spin analysis. The most promising observables are obtained from azimuthal angle correlations in the transverse rest frames of $t(\bar{t})$. This method is particularly useful for determining the coupling structure of $H^\pm$ in the large $\tanβ$ limit, where differences from the SM are most significant.

hep-ph↗

Associated charged Higgs and W boson production in the MSSM at the LHC

We investigate the associated production of charged Higgs bosons (H^\pm) and W bosons at the CERN Large Hadron Collider, using the leptonic decay H^+ -> tau^+ nu_tau and hadronic W decay, within different scenarios of the Minimal Supersymmetric Standard Model (MSSM) with both real and complex parameters. Performing a parton level study we show how the irreducible Standard Model background from W + 2 jets can be controlled by applying appropriate cuts. In the standard m_h^max scenario we find a viable signal for large tan beta and intermediate H^\pm masses (~ m_t). In MSSM scenarios with large mass-splittings among the heavy Higgs bosons the cross-section can be resonantly enhanced by factors up to one hundred, with a strong dependence on the CP-violating phases.

hep-ph↗

H^\pm W^\mp production in the MSSM at the LHC

We investigate the viability of observing charged Higgs bosons (H^\pm) produced in association with W bosons at the CERN Large Hadron Collider, using the leptonic decay H^+ -> tau^+ nu_tau and hadronic W decay, within the Minimal Supersymmetric Standard Model. Performing a parton level study we show how the irreducible Standard Model background from W + 2 jets can be controlled by applying appropriate cuts. In the standard m_h^max scenario we find a viable signal for large tan beta and intermediate H^\pm masses (~ m_t).

hep-ph↗