Searcharxiv⌕ Search

arXiv subjects

Johannes O. Royset

Publications and source records attributed to Johannes O. Royset.

At least 19 recordsLinked to original sources

On Stability in Optimistic Bilevel Optimization

Solutions of bilevel optimization problems tend to suffer from instability under changes to problem data. In the optimistic setting, we construct a lifted formulation that exhibits desirable stability properties under mild assumptions that neither invoke convexity nor smoothness. The upper- and lower-level problems might involve integer restrictions and disjunctive constraints. In a range of results, we invoke at most pointwise and local calmness for the lower-level problem in a sense that holds broadly. The lifted formulation is computationally attractive with structural properties being brought out and an outer approximation algorithm becoming available.

math.OC↗

Perturbation Duality for Robust and Distributionally Robust Optimization: Short and General Proofs

Duality is a foundational tool in robust and distributionally robust optimization (RO/DRO), underpinning both analytical insights and tractable reformulations. Whereas RO/DRO duality is commonly established through minimax arguments or conic duality, we use perturbation duality to obtain new, more general results with short proofs. We show that this perspective provides a natural and unifying framework for deriving RO/DRO dual formulations, proving the associated duality results, and diagnosing the regularity assumptions on which they depend. First, guided by perturbation duality, we establish new duality theorems for a recent DRO framework that unifies several canonical models, including $ϕ$-divergence and Wasserstein models, through optimal transport subject to conditional moment constraints. Our results resolve an open conjecture on this DRO duality by clarifying the role of compactness: compactness itself is not necessary, but can be replaced by perturbation-based regularity conditions. Second, we revisit \emph{robust duality}, commonly described as \emph{primal-worst equals dual-best.} Using bifunctions, we unify dual-best formulations appearing in the literature and derive concise perturbation-based proofs that streamline recent results. Overall, the paper positions perturbation duality as a versatile and underutilized tool for RO and DRO, offering both conceptual unification and technical generality across a broad class of models.

math.OC↗

Lipschitzian SLLNs for random functions

We prove strong laws of large numbers for locally Lipschitz functions in the Lipschitz pseudometric. Our results hold under either a topological or a model-theoretic condition, with the latter encompassing functions jointly definable in o-minimal structures but extending substantially beyond this class. Applications include uniform convergence of limiting and Clarke subdifferentials and finite-sample identification of solutions. Consequently, we identify broad classes of functions for which the failure phenomena revealed by our previous negative results [Tian and Royset, arXiv:2511.16568, 2025] do not occur.

math.OC↗

Failure of uniform laws of large numbers for subdifferentials and beyond

We provide counterexamples showing that uniform laws of large numbers do not hold for subdifferentials under natural assumptions. Our constructions are univariate random Lipschitz functions and bivariate random convex functions with two smooth pieces. Consequently, they resolve the questions posed by Shapiro and Xu [J. Math. Anal. Appl., 325 (2007), 1390-1399] in the negative. They also demonstrate the failure of certain graphical and pointwise laws for subdifferentials, revealing fundamental barriers to the consistency of sample-average approximation and subdifferential approximation.

math.OC↗

Composite Optimization using Local Models and Global Approximations

This work presents a unified framework that combines global approximations with locally built models to handle challenging nonconvex and nonsmooth composite optimization problems, including cases involving extended real-valued functions. We show that near-stationary points of the approximating problems converge to stationary points of the original problem under suitable conditions. Building on this, we develop practical algorithms that use tractable convex master programs derived from local models of the approximating problems. The resulting double-loop structure improves global approximations while adapting local models, providing a flexible and implementable approach for a wide class of composite optimization problems. It also lays the groundwork for new algorithmic developments in this domain.

math.OC↗

Optimistic Bilevel Optimization with Composite Lower-Level Problem

This paper introduces a novel double regularization scheme for bilevel optimization problems whose lower-level problem is composite and convex, but not necessarily strongly convex, in the lower-level variable. The analysis focuses on the primal-dual solution mapping of the regularized lower-level problem and exploits its properties to derive an almost-everywhere formula for the gradient of the regularized hyper-objective under mild assumptions. The paper then establishes conditions under which the hyper-objective of the actual problem is well defined and shows that its gradient can be approximated by the gradient of the regularized hyper-objective. Building on these results, a gradient sampling-based algorithm computes approximately stationary points of the regularized hyper-objective, and we prove its convergence to stationary points of the actual problem. Two numerical examples from machine learning demonstrate the proposed approach.

math.OC↗

Membership Privacy Risks of Sharpness Aware Minimization

Optimization algorithms that seek flatter minima, such as Sharpness-Aware Minimization (SAM), are credited with improved generalization and robustness to noise. We ask whether such gains impact membership privacy. Surprisingly, we find that SAM is more prone to Membership Inference Attacks (MIA) than classical SGD across multiple datasets and attack methods, despite achieving lower test error. This suggests that the geometric mechanism of SAM that improves generalization simultaneously exacerbates membership leakage. We investigate this phenomenon through extensive analysis of memorization and influence scores. Our results reveal that SAM is more capable of capturing atypical subpatterns, leading to higher memorization scores of samples. Conversely, SGD depends more heavily on majority features, exhibiting worse generalization on atypical subgroups and lower memorization. Crucially, this characteristic of SAM can be linked to lower variance in the prediction confidence of unseen samples, thereby amplifying membership signals. Finally, we model SAM under a perfectly interpolating linear regime and theoretically show that sharpness regularization inherently reduces variance, guaranteeing a higher MIA advantage for confidence and likelihood ratio attacks.

cs.LG↗

Rockafellian Relaxation for PDE-Constrained Optimization with Distributional Uncertainty

Stochastic optimization problems are generally known to be ill-conditioned to the form of the underlying uncertainty. A framework is introduced for optimal control problems with partial differential equations as constraints that is robust to inaccuracies in the precise form of the problem uncertainty. The framework is based on problem relaxation and involves optimizing a bivariate, "Rockafellian" objective functional that features both a standard control variable and an additional perturbation variable that handles the distributional ambiguity. In the presence of distributional corruption, the Rockafellian objective functionals are shown in the appropriate settings to $Γ$-converge to uncorrupted objective functionals in the limit of vanishing corruption. Numerical examples illustrate the framework's utility for outlier detection and removal and for variance reduction.

math.OC↗

Epi-Consistent Approximation of Stochastic Dynamic Programs

We study the consistency of stochastic dynamic programs under converging probability distributions and other approximations. Utilizing results on the epi-convergence of expectation functions with varying measures and integrands, and the Attouch--Wets distance, we show that appropriate equi-semicontinuity assumptions assure epi-consistency. A number of examples illustrate the approach. In particular, we permit both unbounded and simultaneously approximated stage-cost functions, and treat an example with approximated constraints.

math.OC↗

Approximating Rockafellians Mitigate Distributional Perturbations: Discontinuous Integrands and Chance-Constrained Applications

In this paper, we show how approximating Rockafellians serve as a principled and effective alternative for improving the stability of stochastic programs under distributional changes. Unlike previous efforts that focus on special distributions and continuous integrands, our results accommodate general probability distributions and discontinuous integrands. Thus, our results apply to chance-constrained programs, for which we obtain improved qualitative and quantitative stability results under weaker assumptions pertaining to metric subregularity and upper outer-Minkowski content.

math.OC↗

Approximations of Rockafellians, Lagrangians, and Dual Functions

Solutions of an optimization problem are sensitive to changes caused by approximations or parametric perturbations, especially in the nonconvex setting. This paper shows that solutions of substitute problems, constructed from Rockafellian functions, can be less sensitive to such changes. Unlike classical stability analysis focused on local changes around (local) minimizers, we employ epi-convergence to examine whether approximating or perturbed problems suitably approach an actual (unperturbed) problem globally. \redrevvv{We demonstrate that solutions derived from the Rockafellian-based substitute problems converge to solutions of the actual optimization problem under suitable conditions, providing a rigorous alternative to potentially unstable direct approximations.} We quantify the rates of convergence that often lead to Lipschitz-kind stability properties for the substitute problems.

math.OC↗

Variational Analysis of a Nonconvex and Nonsmooth Optimization Problem: An Introduction

Variational analysis provides the theoretical foundations and practical tools for constructing optimization algorithms without being restricted to smooth or convex problems. We survey the central concepts in the context of a concrete but broadly applicable problem class from composite optimization in finite dimensions. While prioritizing accessibility over mathematical details, we introduce subgradients of arbitrary functions and the resulting optimality conditions, describe approximations and the need for going beyond pointwise and uniform convergence, and summarize proximal methods. We derive dual problems from parametrization of the actual problem and the resulting relaxations. The paper ends with an introduction to second-order theory and its role in stability analysis of optimization problems.

math.OC↗

Mitigating the Impact of Labeling Errors on Training via Rockafellian Relaxation

Labeling errors in datasets are common, arising in a variety of contexts, such as human labeling, noisy labeling, and weak labeling (i.e., image classification). Although neural networks (NNs) can tolerate modest amounts of these errors, their performance degrades substantially once error levels exceed a certain threshold. We propose a new loss reweighting, architecture-independent methodology, Rockafellian Relaxation Method (RRM) for neural network training. Experiments indicate RRM can enhance neural network methods to achieve robust performance across classification tasks in computer vision and natural language processing (sentiment analysis). We find that RRM can mitigate the effects of dataset contamination stemming from both (heavy) labeling error and/or adversarial perturbation, demonstrating effectiveness across a variety of data domains and machine learning tasks.

cs.LG↗

An implementable proximal-type method for computing critical points to minimization problems with a nonsmooth and nonconvex constraint

This work proposes an implementable proximal-type method for a broad class of optimization problems involving nonsmooth and nonconvex objective and constraint functions. In contrast to existing methods that rely on an ad hoc model approximating the nonconvex functions, our approach can work with a nonconvex model constructed by the pointwise minimum of finitely many convex models. The latter can be chosen with reasonable flexibility to better fit the underlying functions' structure. We provide a unifying framework and analysis covering several subclasses of composite optimization problems and show that our method computes points satisfying certain necessary optimality conditions, which we will call model criticality. Depending on the specific model being used, our general concept of criticality boils down to standard necessary optimality conditions. Numerical experiments on some stochastic reliability-based optimization problems illustrate the practical performance of the method.

math.OC↗

A variational approach to a cumulative distribution function estimation problem under stochastic ambiguity

We propose a method for finding a cumulative distribution function (cdf) that minimizes the distance to a given cdf, while belonging to an ambiguity set constructed relative to another cdf and, possibly, incorporating soft information. Our method embeds the family of cdfs onto the space of upper semicontinuous functions endowed with the hypo-distance. In this setting, we present an approximation scheme based on epi-splines, defined as piecewise polynomial functions, and use bounds for estimating the hypo-distance. Under appropriate hypotheses, we guarantee that the cluster points corresponding to the sequence of minimizers of the resulting approximating problems are solutions to a limiting problem. We describe a large class of functions that satisfy these hypotheses. The approximating method produces a linear-programming-based approximation scheme, enabling us to develop an algorithm from off-the-shelf solvers. The convergence of our proposed approximation is illustrated by numerical examples for the bivariate case.

math.OC↗

Risk-Adaptive Local Decision Rules

For parameterized mixed-binary optimization problems, we construct local decision rules that prescribe near-optimal courses of action across a set of parameter values. The decision rules stem from solving risk-adaptive training problems over classes of continuous, possibly nonlinear mappings. In asymptotic and nonasymptotic analysis, we establish that the decision rules prescribe near-optimal decisions locally for the actual problems, without relying on linearity, convexity, or smoothness. The development also accounts for practically important aspects such as inexact function evaluations, solution tolerances in training problems, regularization, and reformulations to solver-friendly models. The decision rules also furnish a means to carry out sensitivity and stability analysis for broad classes of parameterized optimization problems. We develop a decomposition algorithm for solving the resulting training problems and demonstrate its ability to generate quality decision rules on a nonlinear binary optimization model from search theory.

math.OC↗

Risk-Adaptive Approaches to Stochastic Optimization: A Survey

Uncertainty is prevalent in engineering design, data-driven problems, and decision making broadly. Due to inherent risk-averseness and ambiguity about assumptions, it is common to address uncertainty by formulating and solving conservative optimization models expressed using measures of risk and related concepts. We survey the rapid development of risk measures over the last quarter century. From their beginning in financial engineering, we recount the spread to nearly all areas of engineering and applied mathematics. Solidly rooted in convex analysis, risk measures furnish a general framework for handling uncertainty with significant computational and theoretical advantages. We describe the key facts, list several concrete algorithms, and provide an extensive list of references for further reading. The survey recalls connections with utility theory and distributionally robust optimization, points to emerging applications areas such as fair machine learning, and defines measures of reliability.

math.OC↗

Rockafellian Relaxation and Stochastic Optimization under Perturbations

In practice, optimization models are often prone to unavoidable inaccuracies due to dubious assumptions and corrupted data. Traditionally, this placed special emphasis on risk-based and robust formulations, and their focus on ``conservative" decisions. We develop, in contrast, an ``optimistic" framework based on Rockafellian relaxations in which optimization is conducted not only over the original decision space but also jointly with a choice of model perturbation. The framework enables us to address challenging problems with ambiguous probability distributions from the areas of two-stage stochastic optimization without relatively complete recourse, probability functions lacking continuity properties, expectation constraints, and outlier analysis. We are also able to circumvent the fundamental difficulty in stochastic optimization that convergence of distributions fails to guarantee convergence of expectations. The framework centers on the novel concepts of exact and limit-exact Rockafellians, with interpretations of ``negative'' regularization emerging in certain settings. We illustrate the role of Phi-divergence, examine rates of convergence under changing distributions, and explore extensions to first-order optimality conditions. The main development is free of assumptions about convexity, smoothness, and even continuity of objective functions. Numerical results in the setting of computer vision and text analytics with label noise illustrate the framework.

math.OC↗