SearcharxivSearch

arXiv subjects

Jonathan Yu-Meng Li

Publications and source records attributed to Jonathan Yu-Meng Li.

12 recordsLinked to original sources

Generative Distributionally Robust Optimization

Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibility and adversarial structure: methods that accept arbitrary samplers do not restrict worst-case laws to a generator family, while generator-parameterized adversaries rely on model-specific access such as likelihoods, scores, or training data. We propose Generative Distributionally Robust Optimization (GDRO), a principled framework that accepts any sampleable conditional generator as the nominal model and restricts worst-case laws to a chosen conditional generator family. The key is the sampler-Sinkhorn pairing: samplers represent the conditional laws exactly, while Sinkhorn divergence compares their induced distributions without likelihood access and can be estimated from samples alone. The resulting population problem admits a direct finite-sample approximation and differentiable primal-dual implementation at the active decision context. For Lipschitz losses, the population Sinkhorn radius bounds downstream degradation. Across explicit and implicit generators, our method reduces rare-context inventory regret by 60% and SocialGAN navigation collisions by 50% relative to nominal decisions.

cs.LG

Harnessing Heterogeneous Data for Conditional Optimization via Optimal Transport

Conditional optimization tailors decisions to contextual or event information, but its practical use is often limited by the difficulty of learning the relevant conditional distribution from finite samples of a target joint distribution. This challenge is especially acute when target joint data are scarce or unavailable, or when few observations fall in the conditioning region of interest. Related joint data may be available from multiple sources, such as different stores, markets, populations, or operating environments, but these sources may be biased relative to the target distribution and cannot be pooled naively. We develop a distributionally robust framework based on optimal transport (OT) for harnessing such heterogeneous data in conditional optimization. The framework constructs ambiguity sets over joint distributions using OT distances to empirical source distributions and optimizes worst-case conditional performance over plausible target laws. We propose three OT ambiguity sets that capture different ways of using heterogeneous sources: enforcing simultaneous source consistency, aggregating source discrepancies through weights, and centering the ambiguity set at an OT barycenter. We derive tractable reformulations, establish feasibility conditions, discuss parameter choices, and characterize the relationships among the formulations, revealing trade-offs between robustness, information aggregation, and computational complexity. We demonstrate the value of the framework through a conditional assortment problem using demand and product-feature data from multiple stores.

math.OC

Sampler-Robust Optimization under Generative Models

Modern stochastic optimization pipelines increasingly rely on learned generative models to represent uncertainty, while downstream decisions are evaluated almost entirely through Monte Carlo scenarios. This shifts the operational object of uncertainty from an explicit probability law to the sampler induced by the learned generator. Reliability therefore depends on two errors: sampler misspecification and finite-simulation error. We propose Sampler-Robust Optimization (SRO), which optimizes decisions against the worst-case sampler induced by perturbing the learned generator. This sampler-first formulation aligns with simulation-based decision pipelines and admits a sharpness-aware interpretation: it favors decisions whose performance is stable under generator perturbations, rather than merely under the nominal sampler. Under a coverage assumption, we show that the empirical worst-case objective provides a high-probability upper certificate for the true population objective, with finite-simulation error partially absorbed by the robustification used to guard against sampler misspecification. The framework accommodates generative models with or without explicit densities and admits efficient minimax procedures. Portfolio-optimization experiments show that SRO produces more stable decisions and improves out-of-sample performance under distribution shift.

math.OC

The Virtue of Sparsity in Complexity

Sparsity or complexity? In modern high-dimensional asset pricing, these are often viewed as competing principles: richer feature spaces appear to favor complexity, while economic intuition has long favored parsimony. We show that this tension is misplaced. We distinguish capacity sparsity-the dimensionality of the candidate feature space-from factor sparsity-the parsimonious structure of priced risks-and argue that the two are complements: expanding capacity enables the discovery of factor sparsity. Revisiting the benchmark empirical design of Didisheim et al. (2025) and pushing it to higher complexity regimes, we show that nonlinear feature expansions combined with basis pursuit yield portfolios whose out-of-sample performance dominates ridgeless benchmarks beyond a critical complexity threshold. The evidence shows that the gains from complexity arise not from retaining more factors, but from enlarging the space from which a sparse structure of priced risks can be identified. The virtue of complexity in asset pricing operates through factor sparsity.

q-fin.GN

Generative Adversarial Regression (GAR): Learning Conditional Risk Scenarios

We propose Generative Adversarial Regression (GAR), a framework for learning conditional risk scenarios through generators aligned with downstream risk objectives. GAR builds on a regression characterization of conditional risk for elicitable functionals, including quantiles, expectiles, and jointly elicitable pairs. We extend this principle from point prediction to generative modeling by training generators whose policy-induced risk matches that of real data under the same context. To ensure robustness across all policies, GAR adopts a minimax formulation in which an adversarial policy identifies worst-case discrepancies in risk evaluation while the generator adapts to eliminate them. This structure preserves alignment with the risk functional across a broad class of policies rather than a fixed, pre-specified set. We illustrate GAR through a tail-risk instantiation based on jointly elicitable $(\mathrm{VaR}, \mathrm{ES})$ objectives. Experiments on S\&P 500 data show that GAR produces scenarios that better preserve downstream risk than unconditional, econometric, and direct predictive baselines while remaining stable under adversarially selected policies.

stat.ML

Conditional Risk Minimization with Side Information: A Tractable, Universal Optimal Transport Framework

Conditional risk minimization arises in high-stakes decisions where risk must be assessed in light of side information, such as stressed economic conditions, specific customer profiles, or other contextual covariates. Constructing reliable conditional distributions from limited data is notoriously difficult, motivating a series of optimal-transport-based proposals that address this uncertainty in a distributionally robust manner. Yet these approaches remain fragmented, each constrained by its own limitations: some rely on point estimates or restrictive structural assumptions, others apply only to narrow classes of risk measures, and their structural connections are unclear. We introduce a universal framework for distributionally robust conditional risk minimization, built on a novel union-ball formulation in optimal transport. This framework offers three key advantages: interpretability, by subsuming existing methods as special cases and revealing their deep structural links; tractability, by yielding convex reformulations for virtually all major risk functionals studied in the literature; and scalability, by supporting cutting-plane algorithms for large-scale conditional risk problems. Applications to portfolio optimization with rank-dependent expected utility highlight the practical effectiveness of the framework, with conditional models converging to optimal solutions where unconditional ones clearly do not.

stat.ML

Reconciling Risk-Aversion Paradoxes in the Distribution-Free Newsvendor Problem: Scarf's Rule Meets Dual Utility

How should a risk-averse newsvendor order optimally under distributional ambiguity? Attempts to extend Scarf's celebrated distribution-free ordering rule using risk measures have led to conflicting prescriptions: CVaR-based models invariably recommend ordering less as risk aversion increases, while mean-standard deviation models -- paradoxically -- suggest ordering more, particularly when ordering costs are high. We resolve this behavioral paradox through a coherent generalization of Scarf's distribution-free framework, modeling risk aversion via distortion functionals from dual utility theory. Despite the generality of this class, we derive closed-form optimal ordering rules for any coherent risk preference. These rules uncover a consistent behavioral principle: a more risk-averse newsvendor may rationally order more when overstocking is inexpensive (i.e., when the cost-to-price ratio is low), but will always order less when ordering is costly. Our framework offers a more nuanced, managerially intuitive, and behaviorally coherent understanding of risk-averse inventory decisions. It exposes the limitations of non-coherent models, delivers interpretable and easy-to-compute ordering rules grounded in coherent preferences, and unifies prior work under a single, tractable approach. We further extend the results to multi-product settings with arbitrary demand dependencies, showing that optimal order quantities remain separable and can be obtained by solving single-product problems independently.

math.OC

Wasserstein-Kelly Portfolios: A Robust Data-Driven Solution to Optimize Portfolio Growth

We introduce a robust variant of the Kelly portfolio optimization model, called the Wasserstein-Kelly portfolio optimization. Our model, taking a Wasserstein distributionally robust optimization (DRO) formulation, addresses the fundamental issue of estimation error in Kelly portfolio optimization by defining a ``ball" of distributions close to the empirical return distribution using the Wasserstein metric and seeking a robust log-optimal portfolio against the worst-case distribution from the Wasserstein ball. Enhancing the Kelly portfolio using Wasserstein DRO is a natural step to take, given many successful applications of the latter in areas such as machine learning for generating robust data-driven solutions. However, naive application of Wasserstein DRO to the growth-optimal portfolio problem can lead to several issues, which we resolve through careful modelling. Our proposed model is both practically motivated and efficiently solvable as a convex program. Using empirical financial data, our numerical study demonstrates that the Wasserstein-Kelly portfolio can outperform the Kelly portfolio in out-of-sample testing across multiple performance metrics and exhibits greater stability.

q-fin.PM

On Generalization and Regularization via Wasserstein Distributionally Robust Optimization

Wasserstein distributionally robust optimization (DRO) has gained prominence in operations research and machine learning as a powerful method for achieving solutions with favorable out-of-sample performance. Two compelling explanations for its success are the generalization bounds derived from Wasserstein DRO and its equivalence to regularization schemes commonly used in machine learning. However, existing results on generalization bounds and regularization equivalence are largely limited to settings where the Wasserstein ball is of a specific type, and the decision criterion takes certain forms of expected functions. In this paper, we show that generalization bounds and regularization equivalence can be obtained in a significantly broader setting, where the Wasserstein ball is of a general type and the decision criterion accommodates any form, including general risk measures. This not only addresses important machine learning and operations management applications but also expands to general decision-theoretical frameworks previously unaddressed by Wasserstein DRO. Our results are strong in that the generalization bounds do not suffer from the curse of dimensionality and the equivalency to regularization is exact. As a by-product, we show that Wasserstein DRO coincides with the recent max-sliced Wasserstein DRO for {\it any} decision criterion under affine decision rules -- resulting in both being efficiently solvable as convex programs via our general regularization results. These general assurances provide a strong foundation for expanding the application of Wasserstein DRO across diverse domains of data-driven decision problems.

cs.LG

Inverse Optimization of Convex Risk Functions

The theory of convex risk functions has now been well established as the basis for identifying the families of risk functions that should be used in risk averse optimization problems. Despite its theoretical appeal, the implementation of a convex risk function remains difficult, as there is little guidance regarding how a convex risk function should be chosen so that it also well represents one's own risk preferences. In this paper, we address this issue through the lens of inverse optimization. Specifically, given solution data from some (forward) risk-averse optimization problems we develop an inverse optimization framework that generates a risk function that renders the solutions optimal for the forward problems. The framework incorporates the well-known properties of convex risk functions, namely, monotonicity, convexity, translation invariance, and law invariance, as the general information about candidate risk functions, and also the feedbacks from individuals, which include an initial estimate of the risk function and pairwise comparisons among random losses, as the more specific information. Our framework is particularly novel in that unlike classical inverse optimization, no parametric assumption is made about the risk function, i.e. it is non-parametric. We show how the resulting inverse optimization problems can be reformulated as convex programs and are polynomially solvable if the corresponding forward problems are polynomially solvable. We illustrate the imputed risk functions in a portfolio selection problem and demonstrate their practical value using real-life data.

math.OC

A General Wasserstein Framework for Data-driven Distributionally Robust Optimization: Tractability and Applications

Data-driven distributionally robust optimization is a recently emerging paradigm aimed at finding a solution that is driven by sample data but is protected against sampling errors. An increasingly popular approach, known as Wasserstein distributionally robust optimization (DRO), achieves this by applying the Wasserstein metric to construct a ball centred at the empirical distribution and finding a solution that performs well against the most adversarial distribution from the ball. In this paper, we present a general framework for studying different choices of a Wasserstein metric and point out the limitation of the existing choices. In particular, while choosing a Wasserstein metric of a higher order is desirable from a data-driven perspective, given its less conservative nature, such a choice comes with a high price from a robustness perspective - it is no longer applicable to many heavy-tailed distributions of practical concern. We show that this seemingly inevitable trade-off can be resolved by our framework, where a new class of Wasserstein metrics, called coherent Wasserstein metrics, is introduced. Like Wasserstein DRO, distributionally robust optimization using the coherent Wasserstein metrics, termed generalized Wasserstein distributionally robust optimization (GW-DRO), has all the desirable performance guarantees: finite-sample guarantee, asymptotic consistency, and computational tractability. The worst-case expectation problem in GW-DRO is in general a nonconvex optimization problem, yet we provide new analysis to prove its tractability without relying on the common duality scheme. Our framework, as shown in this paper, offers a fruitful opportunity to design novel Wasserstein DRO models that can be applied in various contexts such as operations management, finance, and machine learning.

math.OC

Closed-form solutions for worst-case law invariant risk measures with application to robust portfolio optimization

Worst-case risk measures refer to the calculation of the largest value for risk measures when only partial information of the underlying distribution is available. For the popular risk measures such as Value-at-Risk (VaR) and Conditional Value-at-Risk (CVaR), it is now known that their worst-case counterparts can be evaluated in closed form when only the first two moments are known for the underlying distribution. These results are remarkable since they not only simplify the use of worst-case risk measures but also provide great insight into the connection between the worst-case risk measures and existing risk measures. We show in this paper that somewhat surprisingly similar closed-form solutions also exist for the general class of law invariant coherent risk measures, which consists of spectral risk measures as special cases that are arguably the most important extensions of CVaR. We shed light on the one-to-one correspondence between a worst-case law invariant risk measure and a worst-case CVaR (and a worst-case VaR), which enables one to carry over the development of worst-case VaR in the context of portfolio optimization to the worst-case law invariant risk measures immediately.

q-fin.RM