SearcharxivSearch

arXiv subjects

Michael McCourt

Publications and source records attributed to Michael McCourt.

At least 19 recordsLinked to original sources

Achieving Diversity in Objective Space for Sample-efficient Search of Multiobjective Optimization Problems

Efficiently solving multi-objective optimization problems for simulation optimization of important scientific and engineering applications such as materials design is becoming an increasingly important research topic. This is due largely to the expensive costs associated with said applications, and the resulting need for sample-efficient, multiobjective optimization methods that efficiently explore the Pareto frontier to expose a promising set of design solutions. We propose moving away from using explicit optimization to identify the Pareto frontier and instead suggest searching for a diverse set of outcomes that satisfy user-specified performance criteria. This method presents decision makers with a robust pool of promising design decisions and helps them better understand the space of good solutions. To achieve this outcome, we introduce the Likelihood of Metric Satisfaction (LMS) acquisition function, analyze its behavior and properties, and demonstrate its viability on various problems.

cs.AI

A Validated Nonlinear Kelvin-Helmholtz Benchmark for Numerical Hydrodynamics

The nonlinear evolution of the Kelvin-Helmholtz instability is a popular test for code verification. To date, most Kelvin-Helmholtz problems discussed in the literature are ill-posed: they do not converge to any single solution with increasing resolution. This precludes comparisons among different codes and severely limits the utility of the Kelvin-Helmholtz instability as a test problem. The lack of a reference solution has led various authors to assert the accuracy of their simulations based on ad-hoc proxies, e.g., the existence of small-scale structures. This paper proposes well-posed Kelvin-Helmholtz problems with smooth initial conditions and explicit diffusion. We show that in many cases numerical errors/noise can seed spurious small-scale structure in Kelvin-Helmholtz problems. We demonstrate convergence to a reference solution using both Athena, a Godunov code, and Dedalus, a pseudo-spectral code. Problems with constant initial density throughout the domain are relatively straightforward for both codes. However, problems with an initial density jump (which are the norm in astrophysical systems) exhibit rich behavior and are more computationally challenging. In the latter case, Athena simulations are prone to an instability of the inner rolled-up vortex; this instability is seeded by grid-scale errors introduced by the algorithm, and disappears as resolution increases. Both Athena and Dedalus exhibit late-time chaos. Inviscid simulations are riddled with extremely vigorous secondary instabilities which induce more mixing than simulations with explicit diffusion. Our results highlight the importance of running well-posed test problems with demonstrated convergence to a reference solution. To facilitate future comparisons, we include the resolved, converged solutions to the Kelvin-Helmholtz problems in this paper in machine-readable form.

astro-ph.IM

Bayesian Optimization is Superior to Random Search for Machine Learning Hyperparameter Tuning: Analysis of the Black-Box Optimization Challenge 2020

This paper presents the results and insights from the black-box optimization (BBO) challenge at NeurIPS 2020 which ran from July-October, 2020. The challenge emphasized the importance of evaluating derivative-free optimizers for tuning the hyperparameters of machine learning models. This was the first black-box optimization challenge with a machine learning emphasis. It was based on tuning (validation set) performance of standard machine learning models on real datasets. This competition has widespread impact as black-box optimization (e.g., Bayesian optimization) is relevant for hyperparameter tuning in almost every machine learning project as well as many applications outside of machine learning. The final leaderboard was determined using the optimization performance on held-out (hidden) objective functions, where the optimizers ran without human intervention. Baselines were set using the default settings of several open-source black-box optimization packages as well as random search.

cs.LG

Bayesian Optimization with Approximate Set Kernels

We propose a practical Bayesian optimization method over sets, to minimize a black-box function that takes a set as a single input. Because set inputs are permutation-invariant, traditional Gaussian process-based Bayesian optimization strategies which assume vector inputs can fall short. To address this, we develop a Bayesian optimization method with \emph{set kernel} that is used to build surrogate functions. This kernel accumulates similarity over set elements to enforce permutation-invariance, but this comes at a greater computational cost. To reduce this burden, we propose two key components: (i) a more efficient approximate set kernel which is still positive-definite and is an unbiased estimator of the true set kernel with upper-bounded variance in terms of the number of subsamples, (ii) a constrained acquisition function optimization over sets, which uses symmetry of the feasible region that defines a set input. Finally, we present several numerical experiments which demonstrate that our method outperforms other methods.

stat.ML

Efficient Rollout Strategies for Bayesian Optimization

Bayesian optimization (BO) is a class of sample-efficient global optimization methods, where a probabilistic model conditioned on previous observations is used to determine future evaluations via the optimization of an acquisition function. Most acquisition functions are myopic, meaning that they only consider the impact of the next function evaluation. Non-myopic acquisition functions consider the impact of the next $h$ function evaluations and are typically computed through rollout, in which $h$ steps of BO are simulated. These rollout acquisition functions are defined as $h$-dimensional integrals, and are expensive to compute and optimize. We show that a combination of quasi-Monte Carlo, common random numbers, and control variates significantly reduce the computational burden of rollout. We then formulate a policy-search based approach that removes the need to optimize the rollout acquisition function. Finally, we discuss the qualitative behavior of rollout policies in the setting of multi-modal objectives and model error.

cs.LG

Sampling Humans for Optimizing Preferences in Coloring Artwork

Many circumstances of practical importance have performance or success metrics which exist implicitly---in the eye of the beholder, so to speak. Tuning aspects of such problems requires working without defined metrics and only considering pairwise comparisons or rankings. In this paper, we review an existing Bayesian optimization strategy for determining most-preferred outcomes, and identify an adaptation to allow it to handle ties. We then discuss some of the issues we have encountered when humans use this optimization strategy to optimize coloring a piece of abstract artwork. We hope that, by participating in this workshop, we can learn how other researchers encounter difficulties unique to working with humans in the loop.

stat.ML

Orchestrate: Infrastructure for Enabling Parallelism during Hyperparameter Optimization

Two key factors dominate the development of effective production grade machine learning models. First, it requires a local software implementation and iteration process. Second, it requires distributed infrastructure to efficiently conduct training and hyperparameter optimization. While modern machine learning frameworks are very effective at the former, practitioners are often left building ad hoc frameworks for the latter. We present SigOpt Orchestrate, a library for such simultaneous training in a cloud environment. We describe the motivating factors and resulting design of this library, feedback from initial testing, and future goals.

cs.DC

A Nonstationary Designer Space-Time Kernel

In spatial statistics, kriging models are often designed using a stationary covariance structure; this translation-invariance produces models which have numerous favorable properties. This assumption can be limiting, though, in circumstances where the dynamics of the model have a fundamental asymmetry, such as in modeling phenomena that evolve over time from a fixed initial profile. We propose a new nonstationary kernel which is only defined over the half-line to incorporate time more naturally in the modeling process.

stat.CO

On the Dynamics of the Inclination Instability

Axisymmetric disks of eccentric Kepler orbits are vulnerable to an instability which causes orbits to exponentially grow in inclination, decrease in eccentricity, and cluster in their angle of pericenter. Geometrically, the disk expands to a cone shape which is asymmetric about the mid-plane. In this paper, we describe how secular gravitational torques between individual orbits drive this "inclination instability". We derive growth timescales for a simple two-orbit model using a Gauss $N$-ring code, and generalize our result to larger $N$ systems with $N$-body simulations. We find that two-body relaxation slows the growth of the instability at low $N$ and that angular phase coverage of orbits in the disk is important at higher $N$. As $N \to \infty$, the e-folding timescale converges to that expected from secular theory.

astro-ph.EP

The impact of magnetic fields on thermal instability

Cold ($T\sim 10^{4} \ \mathrm{K}$) gas is very commonly found in both galactic and cluster halos. There is no clear consensus on its origin. Such gas could be uplifted from the central galaxy by galactic or AGN winds. Alternatively, it could form in situ by thermal instability. Fragmentation into a multi-phase medium has previously been shown in hydrodynamic simulations to take place once $t_\mathrm{cool}/t_\mathrm{ff}$, the ratio of the cooling time to the free-fall time, falls below a threshold value. Here, we use 3D plane-parallel MHD simulations to investigate the influence of magnetic fields. We find that because magnetic tension suppresses buoyant oscillations of condensing gas, it destabilizes all scales below $l_\mathrm{A}^\mathrm{cool} \sim v_\mathrm{A} t_\mathrm{cool}$, enhancing thermal instability. This effect is surprisingly independent of magnetic field orientation or cooling curve shape, and sets in even at very low magnetic field strengths. Magnetic fields critically modify both the amplitude and morphology of thermal instability, with $δρ/ρ\propto β^{-1/2}$, where $β$ is the ratio of thermal to magnetic pressure. In galactic halos, magnetic fields can render gas throughout the entire halo thermally unstable, and may be an attractive explanation for the ubiquity of cold gas, even in the halos of passive, quenched galaxies.

astro-ph.GA

Sequential Preference-Based Optimization

Many real-world engineering problems rely on human preferences to guide their design and optimization. We present PrefOpt, an open source package to simplify sequential optimization tasks that incorporate human preference feedback. Our approach extends an existing latent variable model for binary preferences to allow for observations of equivalent preference from users.

cs.LG

Practical Bayesian optimization in the presence of outliers

Inference in the presence of outliers is an important field of research as outliers are ubiquitous and may arise across a variety of problems and domains. Bayesian optimization is method that heavily relies on probabilistic inference. This allows outstanding sample efficiency because the probabilistic machinery provides a memory of the whole optimization process. However, that virtue becomes a disadvantage when the memory is populated with outliers, inducing bias in the estimation. In this paper, we present an empirical evaluation of Bayesian optimization methods in the presence of outliers. The empirical evidence shows that Bayesian optimization with robust regression often produces suboptimal results. We then propose a new algorithm which combines robust regression (a Gaussian process with Student-t likelihood) with outlier diagnostics to classify data points as outliers or inliers. By using an scheduler for the classification of outliers, our method is more efficient and has better convergence over the standard robust regression. Furthermore, we show that even in controlled situations with no expected outliers, our method is able to produce better results.

cs.LG

Active Preference Learning for Personalized Portfolio Construction

In financial asset management, choosing a portfolio requires balancing returns, risk, exposure, liquidity, volatility and other factors. These concerns are difficult to compare explicitly, with many asset managers using an intuitive or implicit sense of their interaction. We propose a mechanism for learning someone's sense of distinctness between portfolios with the goal of being able to identify portfolios which are predicted to perform well but are distinct from the perspective of the user. This identification occurs, e.g., in the context of Bayesian optimization of a backtested performance metric. Numerical experiments are presented which show the impact of personal beliefs in informing the development of a diverse and high-performing portfolio.

cs.CE

Robust Bayesian Optimization with Student-t Likelihood

Bayesian optimization has recently attracted the attention of the automatic machine learning community for its excellent results in hyperparameter tuning. BO is characterized by the sample efficiency with which it can optimize expensive black-box functions. The efficiency is achieved in a similar fashion to the learning to learn methods: surrogate models (typically in the form of Gaussian processes) learn the target function and perform intelligent sampling. This surrogate model can be applied even in the presence of noise; however, as with most regression methods, it is very sensitive to outlier data. This can result in erroneous predictions and, in the case of BO, biased and inefficient exploration. In this work, we present a GP model that is robust to outliers which uses a Student-t likelihood to segregate outliers and robustly conduct Bayesian optimization. We present numerical results evaluating the proposed method in both artificial functions and real problems.

cs.LG

Resonant line transfer in a fog: Using Lyman-alpha to probe tiny structures in atomic gas

Motivated by observational and theoretical work which both suggest very small scale ($\lesssim 1\,$pc) structure in the circum-galactic medium of galaxies and in other environments, we study Lyman-$α$ (Ly$α$) radiative transfer in an extremely clumpy medium with many "clouds" of neutral gas along the line of sight. While previous studies have typically considered radiative transfer through sightlines intercepting $\lesssim 10$ clumps, we explore the limit of a very large number of clumps per sightline (up to $f_{\mathrm{c}} \sim 1000$). Our main finding is that, for covering factors greater than some critical threshold, a multiphase medium behaves similar to a homogeneous medium in terms of the emergent Ly$α$ spectrum. The value of this threshold depends on both the clump column density and on the movement of the clumps. We estimate this threshold analytically and compare our findings to radiative transfer simulations with a range of covering factors, clump column densities, radii, and motions. Our results suggest that (i) the success in fitting observed Ly$α$ spectra using homogeneous "shell models" (and the corresponding failure of multiphase models) hints towards the presence of very small-scale structure in neutral gas, in agreement within a number of other observations; and (ii) the recurrent problems of reproducing realistic line profiles from hydrodynamical simulations may be due to their inability to resolve small-scale structure, which causes simulations to underestimate the effective covering factor of neutral gas clouds.

astro-ph.GA

The Impact of Star Formation Feedback on the Circumgalactic Medium

We use idealized 3D hydrodynamic simulations to study the dynamics and thermal structure of the circumgalactic medium (CGM). Our simulations quantify the role of cooling, stellar feedback driven galactic winds and cosmological gas accretion in setting the properties of the CGM in dark matter haloes ranging from $10^{11}$ to $10^{12}$ M$_\odot$. Our simulations support a conceptual picture in which the properties of the CGM, and the key physics governing it, change markedly near a critical halo mass of M$_{\rm crit} \approx 10^{11.5}$ M$_\odot$. As in calculations without stellar feedback, above M$_{\rm crit}$ halo gas is supported by thermal pressure created in the virial shock. The thermal properties at small radii are regulated by feedback triggered when $t_{\rm cool}/t_{\rm ff}\lesssim10$ in the hot gas. Below M$_{\rm crit}$, however, there is no thermally supported halo and self-regulation at $t_{\rm cool}/t_{\rm ff}\sim10$ does not apply. Instead, the gas is out of hydrostatic equilibrium and largely supported against gravity by bulk flows (turbulence and coherent inflow/outflow) arising from the interaction between cosmological gas inflow and outflowing galactic winds. In these lower mass haloes, the phase structure depends sensitively on the outflows' energy per unit mass and mass-loading, which may allow measurements of the CGM thermal state to constrain the nature of galactic winds. Our simulations account for some of the properties of the multiphase halo gas inferred from quasar absorption line observations, including the presence of significant mass at a wide range of temperatures, and the characteristic OVI and CIV column densities and kinematics. However, we underpredict the neutral hydrogen content of the $z\sim0$ CGM.

astro-ph.GA

Simulations of Magnetic Fields in Tidally-Disrupted Stars

We perform the first magnetohydrodynamical simulations of tidal disruptions of stars by supermassive black holes. We consider stars with both tangled and ordered magnetic fields, for both grazing and deeply disruptive encounters. When the star survives disruption, we find its magnetic field amplifies by a factor of up to twenty, but see no evidence for the a self-sustaining dynamo that would yield arbitrary field growth. For stars that do not survive, and within the tidal debris streams produced in partial disruptions, we find that the component of the magnetic field parallel to the direction of stretching along the debris stream only decreases slightly with time, eventually resulting in a stream where the magnetic pressure is in equipartition with the gas. Our results suggest that the returning gas in most (if not all) stellar tidal disruptions is already highly magnetized by the time it returns to the black hole.

astro-ph.HE

Interactive Preference Learning of Utility Functions for Multi-Objective Optimization

Real-world engineering systems are typically compared and contrasted using multiple metrics. For practical machine learning systems, performance tuning is often more nuanced than minimizing a single expected loss objective, and it may be more realistically discussed as a multi-objective optimization problem. We propose a novel generative model for scalar-valued utility functions to capture human preferences in a multi-objective optimization setting. We also outline an interactive active learning system that sequentially refines the understanding of stakeholders ideal utility functions using binary preference queries.

math.OC