Searcharxiv⌕ Search

arXiv · 2609.32772

Learnable Randomization as Commitment Against Adaptive Optimizers

Abstract

A pricing page can walk the posted price up to the last amount a buyer still accepts, a recommender can hold back a better item for a barely acceptable promoted one, and a classifier can shift its boundary once applicants change their features. The system predicts the response and then picks the menu that serves its own objective, so the surplus above the user's cutoff is taken. Playing the single best action publishes that cutoff, while noise on actions the user would never take throws away payoff and teaches the platform that a worse menu is still acceptable. We study unpredictable near-optimal policies (UNOP), which mix uniformly on near-best actions that remain individually rational. The mixture is a commitment about the response. On a finite price grid, when the best sure-demand price strictly out-earns the randomized band, a seller who already knows the curve posts below the band, and the purchase that occurs is deterministic. Knowing that curve is not the same as predicting the next draw. The mixture can be learned and the optimizer can match its best response, while the user's payoff stays higher because the mixture changes which action is targeted. In pricing and in policy-aware recommendation this leaves more surplus than greedy play when the platform optimizes against the curve and more than one action is acceptable. The gain goes away under quality ranking, a singleton near-optimal set, a wrong utility estimate, or a short-horizon explorer. That is also where mixing should be turned off if the other side is trying to cooperate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zihan Deng, Chuanzhi Xu, Xiaozhen Zhong, Haoyang Li, Junjie Huang. 2026-09-26. Learnable Randomization as Commitment Against Adaptive Optimizers. https://arxiv.org/abs/2609.32772

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders

Training on vast amounts of human-generated data has motivated growing interest in using large language models (LLMs) to simulate human behavior. We ask which features of human behavior general-purpose models preserve when used out of the box in auctions, where multiple bidders interact under explicit rules and incentives. We evaluate five LLMs across seven laboratory settings against human benchmarks reconstructed from published experiments, with uncertainty bands for the private-value comparisons. Our main focus is on three large models without extended test-time reasoning: GPT-4o, Claude~3.5 Haiku, and Gemini~2.0 Flash. LLM and human deviations from theory differ in magnitude and often in direction: humans overbid in second-price auctions, whereas most models that deviate underbid. Surprisingly, without task-specific fine-tuning or calibration to human bids, the three non-reasoning large models robustly preserve key orderings of auction formats by deviation from theory. First-price auctions are harder than second-price, and ascending clocks reduce deviations relative to sealed bids wherever data are adequate. Kendall's $τ_b$ between the human and GPT-4o difficulty rankings is $0.60$ and positive in every joint bootstrap draw. The reasoning model bids almost at equilibrium in the observed private-value settings, leaving little variation in errors to compare; the small model's large errors yield an inverted ranking. All five models nevertheless reproduce the stronger first-price winner's curse. Clock framing improves bidding for two of the three non-reasoning large models, and GPT-4o recovers the ordering of last-minute bidding across closing rules in an eBay-style marketplace.

cs.GT↗

Dynamic Transaction Scheduling and Pricing in the Ethereum Mempool

The Ethereum blockchain utilizes the EIP-1559 algorithm to manage transaction inclusion and block assembly. However, EIP-1559 and much of the existing literature study this problem from a static perspective, focusing on price evolution without modelling transaction dynamics within the mempool. Motivated by this limitation, we study a dynamic transaction scheduling problem in which transactions with heterogeneous sizes and per-unit values arrive over time and remain in the mempool until scheduled. To capture the stochastic mempool evolution, we formulate the problem as a Markov Decision Process (MDP) whose state represents the mempool configuration and whose actions correspond to block prices. We first provide a primal-dual interpretation of the static EIP-1559 mechanism, showing that block prices arise naturally as dual variables of a social-welfare maximization problem. Building on this perspective, we extend the framework to the dynamic setting and formulate an objective that maximizes long-run discounted reward while incorporating holding costs and overshoot penalties. We then employ a Natural Policy Gradient (NPG) algorithm to compute the optimal policy. Our results show that dynamic pricing stabilizes the mempool while maximizing long-run discounted reward. In particular, as the overshoot penalty increases, the average scheduled transaction volume converges to the target block capacity, and the resulting NPG updates closely resemble the EIP-1559 price update rule. Finally, we study two special cases of the MDP formulation: homogeneous transactions and uniform arrivals. In the homogeneous setting, where the protocol directly controls scheduled volume, we show that the optimal policy has a threshold structure. We then propose a bang-bang pricing mechanism for uniform arrivals and derive a lower bound on the block capacity needed to ensure system stability.

cs.GT↗

Computational Complexity of Strong and Average Justified Representation

We study the approval-based multiwinner election problem where a set of $n$ voters cast approval-based ballots to a set of $m$ candidates, and we are to select a winner committee consisting of $k$ candidates. We consider two axioms: strong justified representation (SJR) and average justified representation (AJR). A winner committee satisfies SJR if the satisfaction for each voter in every $\ell$-cohesive group is at least $\ell$. AJR is a weaker axiom that requires the average satisfaction for each $\ell$-cohesive group to be at least $\ell$. It is well known that a winner committee satisfying AJR may not exist (and neither does SJR). In this paper, we study the computational complexity of the following decision problem: given an approval-based multiwinner election instance, decide if there exists a winner committee satisfying SJR/AJR. We prove that this problem is $Θ_2^p$-complete for SJR, and $Σ_2^p$-complete for AJR. As byproducts, we derive some results that are interesting in their own right. Firstly, we show that adding one more adaptive query to an NP oracle on top of polynomially many non-adaptive NP queries does not add more computational power, and the resulting complexity class is still $Θ_2^p$. Secondly, we construct a set system that can be useful in other applications, especially when doing reductions from typical satisfiability problems such as 3SAT.

cs.GT↗