SearcharxivSearch

arXiv subjects

Sara Shashaani

Publications and source records attributed to Sara Shashaani.

13 recordsLinked to original sources

Adaptive Sampling Trust Region Optimization for Derivative-free Stochastic Functions and Deterministic Equality Constraints

We study optimization problems with noisy zeroth-order objective observations and deterministic nonlinear equality constraints with available derivatives. We propose a constrained variant of the adaptive-sampling trust-region derivative-free optimization algorithm---ASTRO-DF. The method builds quadratic local models from estimated objective values at interpolation points within a moving trust region and promotes feasibility through a Byrd--Omojokun composite-step based on linearized constraints, following an SQP-like framework. We prove almost sure convergence using a new constrained criticality test and present numerical results on an equality-constrained stochastic activity network problem.

math.OC

Extrapolation-based Direct Search for Nonsmooth Stochastic Zeroth-Order Optimization

We propose and analyze a stochastic direct-search method for unconstrained zeroth-order minimization of locally Lipschitz, possibly nonsmooth, objectives. The method combines random polling directions with a stochastic extrapolating line search based on a sufficient-decrease test of order $p$. Under conditional accuracy assumptions on the stochastic estimates, we prove almost-sure convergence to Clarke stationary points. We further establish an expected iteration complexity bound. Specifically, using a supermartingale stopping-time argument, we prove that $\mathcal O\left( \max\left\{ r^{-p}, \varepsilon^{-p/(p-1)} \right\} \right) $ iterations are sufficient in expectation to reach an $(r,\varepsilon)$-Goldstein stationary point. Moreover, we derive a corresponding expected tested-point complexity bound of order $\mathcal O\bigl(\varepsilon^{1-n} \max\{r^{-p},\varepsilon^{-p/(p-1)}\}\bigr)$. To the best of our knowledge, this is the first convergence and expected-complexity analysis for an extrapolation-based direct-search method in a nonsmooth stochastic setting. Numerical experiments on a DFO benchmark suite highlight competitive performance against well-established stochastic direct-search methods.

math.OC

Adaptive Regularization within Trust Region Methods for Stochastic Nonconvex Optimization

We propose a stochastic nonconvex optimization algorithm that achieves almost sure $\tilde{\mathcal{O}}(ε^{-1.5})$ iteration complexity for problems with smooth objective functions and gradients only observable with noise. The mean-zero stochastic noise is decision-dependent and has unbounded support with subexponential tail, allowing our framework to cover a broad class of problems. The improved almost sure iteration complexity is achieved with a new variant of the adaptive sampling trust-region optimization (ASTRO) augmented with an adaptively regularized local model, which we term Reg-ASTRO. Adaptive sampling ensures that the estimation precision is aligned with a measure of stationarity, so that iterates closer to stationarity trigger higher accuracy requirement for sampling. A key analytical challenge arises because the trust-region radius and regularization are coupled and not determined prior to gradient estimation at each iteration. We further establish an almost sure $\tilde{\mathcal{O}}(ε^{-4.5})$ sample complexity for Reg-ASTRO, which improves to $\tilde{\mathcal{O}}(ε^{-3.5})$ under stronger regularity conditions and use of common random numbers, substantially outperforming first-order methods in theory and numerical experiments.

math.OC

Stratified adaptive sampling for derivative-free stochastic trust-region optimization

There is emerging evidence that trust-region (TR) algorithms are very effective at solving derivative-free nonconvex stochastic optimization problems in which the objective function is a Monte Carlo (MC) estimate. A recent strand of methodologies adaptively adjusts the sample size of the MC estimates by keeping the estimation error below a measure of stationarity induced from the TR radius. In this work we explore stratified adaptive sampling strategies to equip the TR framework with accurate estimates of the objective function, thus optimizing the required number of MC samples to reach a given ε-accuracy of the solution. We prove a reduced sample complexity, confirm a superior efficiency via numerical tests and applications, and explore inexpensive implementations in high dimension.

math.OC

Root Finding and Metamodeling for Rapid and Robust Computer Model Calibration

We concern computer model calibration problem where the goal is to find the parameters that minimize the discrepancy between the multivariate real-world and computer model outputs. We propose to solve an approximation using signed residuals that enables a root finding approach and an accelerated search. We characterize the distance of the solutions to the approximation from the solutions of the original problem for the strongly-convex objective functions, showing that it depends on variability of the signed residuals across output dimensions, as wells as their variance and covariance. We develop a metamodel-based root finding framework under kriging and stochastic kriging that is augmented with a sequential search space reduction. We derive three new acquisition functions for finding roots of the approximate problem along with their derivatives usable by first-order solvers. Compared to kriging, stochastic kriging accounts for observational noise, promoting more robust solutions. We also analyze the case where a root may not exist. Our analysis of the asymptotic behavior in this context show that, since existence of roots in the approximation problem may not be known a priori, using new acquisition functions will not compromise the outcome. Numerical experiments on data-driven and physics-based examples demonstrate significant computational gains over standard calibration approaches.

stat.ME

Complexity of Zeroth- and First-order Stochastic Trust-Region Algorithms

Model update (MU) and candidate evaluation (CE) are classical steps incorporated inside many stochastic trust-region (TR) algorithms. The sampling effort exerted within these steps, often decided with the aim of controlling model error, largely determines a stochastic TR algorithm's sample complexity. Given that MU and CE are amenable to variance reduction, we investigate the effect of incorporating common random numbers (CRN) within MU and CE on complexity. Using ASTRO and ASTRO-DF as prototype first-order and zeroth-order families of algorithms, we demonstrate that CRN's effectiveness leads to a range of complexities depending on sample-path regularity and the oracle order. For instance, we find that in first-order oracle settings with smooth sample paths, CRN's effect is pronounced -- ASTRO with CRN achieves $\tilde{O}(ε^{-2})$ a.s. sample complexity compared to $\tilde{O}(ε^{-6})$ a.s. in the generic no-CRN setting. By contrast, CRN's effect is muted when the sample paths are not Lipschitz, with the sample complexity improving from $\tilde{O}(ε^{-6})$ a.s. to $\tilde{O}(ε^{-5})$ and $\tilde{O}(ε^{-4})$ a.s. in the zeroth- and first-order settings, respectively. Since our results imply that improvements in complexity are largely inherited from generic aspects of variance reduction, e.g., finite-differencing for zeroth-order settings and sample-path smoothness for first-order settings within MU, we anticipate similar trends in other contexts.

math.OC

Uncertainty Quantification using Simulation Output: Batching as an Inferential Device

We present batching as an omnibus device for uncertainty quantification using simulation output. We consider the classical context of a simulationist performing uncertainty quantification on an estimator $θ_n$ (of an unknown fixed quantity $θ$) using only the output data $(Y_1,Y_2,\ldots,Y_n)$ gathered from a simulation. By uncertainty quantification, we mean approximating the sampling distribution of the error $θ_n-θ$ toward: (A) estimating an assessment functional $ψ$, e.g., bias, variance, or quantile; or (B) constructing a $(1-α)$-confidence region on $θ$. We argue that batching is a remarkably simple and effective device for this purpose, and is especially suited for handling dependent output data such as what one frequently encounters in simulation contexts. We demonstrate that if the number of batches and the extent of their overlap are chosen appropriately, batching retains bootstrap's attractive theoretical properties of strong consistency and higher-order accuracy. For constructing confidence regions, we characterize two limiting distributions associated with a Studentized statistic. Our extensive numerical experience confirms theoretical insight, especially about the effects of batch size and batch overlap.

stat.ME

Simulation Model Calibration with Dynamic Stratification and Adaptive Sampling

Calibrating simulation models that take large quantities of multi-dimensional data as input is a hard simulation optimization problem. Existing adaptive sampling strategies offer a methodological solution. However, they may not sufficiently reduce the computational cost for estimation and solution algorithm's progress within a limited budget due to extreme noise levels and heteroskedasticity of system responses. We propose integrating stratification with adaptive sampling for the purpose of efficiency in optimization. Stratification can exploit local dependence in the simulation inputs and outputs. Yet, the state-of-the-art does not provide a full capability to adaptively stratify the data as different solution alternatives are evaluated. We devise two procedures for data-driven calibration problems that involve a large dataset with multiple covariates to calibrate models within a fixed overall simulation budget. The first approach dynamically stratifies the input data using binary trees, while the second approach uses closed-form solutions based on linearity assumptions between the objective function and concomitant variables. We find that dynamical adjustment of stratification structure accelerates optimization and reduces run-to-run variability in generated solutions. Our case study for calibrating a wind power simulation model, widely used in the wind industry, using the proposed stratified adaptive sampling, shows better-calibrated parameters under a limited budget.

stat.ME

Building Trees for Probabilistic Prediction via Scoring Rules

Decision trees built with data remain in widespread use for nonparametric prediction. Predicting probability distributions is preferred over point predictions when uncertainty plays a prominent role in analysis and decision-making. We study modifying a tree to produce nonparametric predictive distributions. We find the standard method for building trees may not result in good predictive distributions and propose changing the splitting criteria for trees to one based on proper scoring rules. Analysis of both simulated data and several real datasets demonstrates that using these new splitting criteria results in trees with improved predictive properties considering the entire predictive distribution.

stat.ME

Iteration Complexity and Finite-Time Efficiency of Adaptive Sampling Trust-Region Methods for Stochastic Derivative-Free Optimization

Adaptive sampling with interpolation-based trust regions or ASTRO-DF is a successful algorithm for stochastic derivative-free optimization with an easy-to-understand-and-implement concept that guarantees almost sure convergence to a first-order critical point. To reduce its dependence on the problem dimension, we present local models with diagonal Hessians constructed on interpolation points based on a coordinate basis. We also leverage the interpolation points in a direct search manner whenever possible to boost ASTRO-DF's performance in a finite time. We prove that the algorithm has a canonical iteration complexity of $\mathcal{O}(ε^{-2})$ almost surely, which is the first guarantee of its kind without placing assumptions on the quality of function estimates or model quality or independence between them. Numerical experimentation reveals the computational advantage of ASTRO-DF with coordinate direct search due to saving and better steps in the early iterations of the search.

math.OC

Two-Stage Estimation and Variance Modeling for Latency-Constrained Variational Quantum Algorithms

The Quantum Approximate Optimization Algorithm (QAOA) has enjoyed increasing attention in noisy intermediate-scale quantum computing due to its application to combinatorial optimization problems. Because combinatorial optimization problems are NP-hard, QAOA could serve as a potential demonstration of quantum advantage in the future. As a hybrid quantum-classical algorithm, the classical component of QAOA resembles a simulation optimization problem, in which the simulation outcomes are attainable only through the quantum computer. The simulation that derives from QAOA exhibits two unique features that can have a substantial impact on the optimization process: (i) the variance of the stochastic objective values typically decreases in proportion to the optimality gap, and (ii) querying samples from a quantum computer introduces an additional latency overhead. In this paper, we introduce a novel stochastic trust-region method, derived from a derivative-free adaptive sampling trust-region optimization (ASTRO-DF) method, intended to efficiently solve the classical optimization problem in QAOA, by explicitly taking into account the two mentioned characteristics. The key idea behind the proposed algorithm involves constructing two separate local models in each iteration: a model of the objective function, and a model of the variance of the objective function. Exploiting the variance model allows us to both restrict the number of communications with the quantum computer, and also helps navigate the nonconvex objective landscapes typical in the QAOA optimization problems. We numerically demonstrate the superiority of our proposed algorithm using the SimOpt library and Qiskit, when we consider a metric of computational burden that explicitly accounts for communication costs.

math.OC

Robust Output Analysis with Monte-Carlo Methodology

In predictive modeling with simulation or machine learning, it is critical to accurately assess the quality of estimated values through output analysis. In recent decades output analysis has become enriched with methods that quantify the impact of input data uncertainty in the model outputs to increase robustness. However, most developments are applicable assuming that the input data adheres to a parametric family of distributions. We propose a unified output analysis framework for simulation and machine learning outputs through the lens of Monte Carlo sampling. This framework provides nonparametric quantification of the variance and bias induced in the outputs with higher-order accuracy. Our new bias-corrected estimation from the model outputs leverages the extension of fast iterative bootstrap sampling and higher-order influence functions. For the scalability of the proposed estimation methods, we devise budget-optimal rules and leverage control variates for variance reduction. Our theoretical and numerical results demonstrate a clear advantage in building more robust confidence intervals from the model outputs with higher coverage probability.

stat.ME

ASTRO-DF: A Class of Adaptive Sampling Trust-Region Algorithms for Derivative-Free Stochastic Optimization

We consider unconstrained optimization problems where only "stochastic" estimates of the objective function are observable as replicates from a Monte Carlo oracle. The Monte Carlo oracle is assumed to provide no direct observations of the function gradient. We present ASTRO-DF --- a class of derivative-free trust-region algorithms, where a stochastic local interpolation model is constructed, optimized, and updated iteratively. Function estimation and model construction within ASTRO-DF is adaptive in the sense that the extent of Monte Carlo sampling is determined by continuously monitoring and balancing metrics of sampling error (or variance) and structural error (or model bias) within ASTRO-DF. Such balancing of errors is designed to ensure that Monte Carlo effort within ASTRO-DF is sensitive to algorithm trajectory, sampling more whenever an iterate is inferred to be close to a critical point and less when far away. We demonstrate the almost-sure convergence of ASTRO-DF's iterates to a first-order critical point when using linear or quadratic stochastic interpolation models. The question of using more complicated models, e.g., regression or stochastic kriging, in combination with adaptive sampling is worth further investigation and will benefit from the methods of proof presented here. We speculate that ASTRO-DF's iterates achieve the canonical Monte Carlo convergence rate, although a proof remains elusive.

math.OC