Searcharxiv⌕ Search

arXiv · 2610.09843

Minimizing Cumulative Envy in Allocating a Sequence of Items

Abstract

We study temporal fair division with indivisible goods that arrive sequentially and must be allocated irrevocably. In contrast to the usual online model, we assume that valuations and future arrivals are known in advance, and ask how unfairness evolves during the process. We introduce \emph{cumulative maximum envy}: the sum, over all rounds, of the maximum pairwise envy at that round. Equivalently, this is the area under the worst-envy curve, and it captures both the magnitude and the duration of envy. For a fixed arrival order, we show that the corresponding decision problem is strongly NP-complete and that minimizing this objective admits no constant-factor approximation unless P = NP, even under identical valuations and even under binary valuations. We complement these hardness results with a dynamic program that gives pseudopolynomial-time solvability for a constant number of agents, polynomial-time algorithms in further restricted settings, and an FPTAS for fixed $n$ under identical integer valuations. We then study a sequencing variant where the algorithm may choose the arrival order. This variant remains NP-complete even for two agents with identical valuations; however, a simple greedy algorithm achieves a $3/2$-approximation for $n=2$ agents, an $n/(n-1)$-approximation for any number of agents, and an additive guarantee depending on the maximum value of any good.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Paul W. Goldberg, Isaac Robinson, Nicholas Teh. 2026-10-07. Minimizing Cumulative Envy in Allocating a Sequence of Items. https://arxiv.org/abs/2610.09843

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stackelberg POMDP: Learning to Lead via Reinforcement Learning

Many real-world domains--including e-commerce platform design, security planning, and multi-agent coordination--feature leader-follower problems where one decision-maker commits to a policy and others react strategically. We develop a reinforcement learning framework for such interactions in sequential environments with partial observations and multiple followers. Followers may adapt through no-regret learning or reinforcement learning, potentially departing from equilibrium behavior. The framework embeds follower adaptation into the leader's environment to construct a single-agent partially observable Markov decision process--the Stackelberg POMDP. For policy-interactive response algorithms, which access the leader's policy through queries, we prove that an optimal policy based only on the leader's game history yields an optimal commitment under the specified response procedure. We use proximal policy optimization with a centralized critic and train contextual meta-followers to respond across leader policies. In indirect mechanism design, mechanisms using buyer messages achieve higher social welfare than optimal standard sequential price mechanisms across all tested type counts, with responses certified as approximate Bayesian coarse correlated equilibria. In platform design, learned display rules increase mean consumer surplus by 8.4% over an optimized fixed price cap while accommodating hidden seller costs. In Atari bilateral trade, meta-learned follower responses support joint learning of visual gameplay and economic decisions; assigning leadership to the seller or buyer shifts transaction prices and payoffs in that agent's favor. Controlled ablations examine how response credit, policy consistency, and reward timing affect learning.

cs.GT↗

Price Competition Under Platform-Mediated Search: A Consider-Then-Choose Framework

We study the problem of predicting price equilibria on e-commerce platforms where sellers compete across multiple attributes (e.g., price, average rating, delivery speed). In these settings, a platform's design choices --- such as its display ranking, filtering tools, and promotional badges --- critically shape customer search and purchase behavior, which in turn determine sellers' equilibrium pricing strategies. Our goal is to develop a tractable framework that allows a platform to anticipate the market impact of its design interventions. We consider a behavioral model --- Consider-then-Choose with Lexicographic Choice (CLC) --- specifically tailored to platform-mediated search. We establish that any local Nash equilibrium admits a sequential-move characterization; this yields a tractable procedure for computation under an interpretable sufficient condition, which we term gradient dominance. We further prove that under gradient dominance, simple, decentralized gradient-based algorithms converge to an equilibrium, providing platforms with a method for simulating market outcomes. Finally, we use our framework to study how platform design affects market outcomes. Our framework applies to any platform in which sellers compete on multiple attributes and customer choice is guided by the platform's interface. In these environments, sellers' pricing strategies must be understood not in isolation, but as a response to the platform's design. Our work provides platform operators with a rigorous toolbox to efficiently evaluate how changes to interface design, information disclosure, and ranking policies can affect competitive outcomes.

cs.GT↗

From Reconnaissance to Response: Quantitative Risk Parameterization and Game Theoretic Containment in Modern Enterprise Attack

Modern Security Operations Centers struggle with delayed manual incident response, enabling adversaries to advance through the Cyber Kill Chain during early stage reconnaissance. While classical game theoretic defense models optimize strategic resource allocation, they rely on static utility matrices that fail to adapt to dynamic telemetry. This paper presents an integrated, metrics driven decision engine that bridges quantitative risk parameterization and continuous automated response time. Common Vulnerability Scoring Systems exploitability parameters are mapped to attacker success probabilities and evaluate defender log distributions via Factor Analysis of Information Risk Monte Carlo simulations. Real time SIEM logs streams are modeled as Poisson process arrival rates, dynamically updating defender posterior threat belief through sequential Bayesian filtering. A closed form threshold is derived by framing the interaction as a dynamic Bayesian Stackelberg game, where the expected unmitigated risk exceeds proactive containment cost. Parameterized against empirical data from the 2023 MGM Resorts and Caesars Entertainment cyber incident, simulation results demonstrate that the engine suppresses transient background noise while triggering automated SOAR network isolation within seconds of adversarial probing. Multi parameter sensitivity analysis confirms that the decision boundary dynamically adjusts to live perimeter vulnerability, offering a control theoretic foundation for sub minute automated threat containment.

cs.GT↗