Searcharxiv⌕ Search

arXiv · 2609.36589

Non-Linear Pricing Restores Tractability for a Data Seller

Abstract

We consider a data seller who designs pricing mechanisms over multiple datasets to maximize revenue from budget-constrained buyers. The seller offers multiple datasets and assigns each a pricing function that maps the quantity purchased to a total payment. The goal is to design these pricing functions to maximize revenue, anticipating that buyers---who trade off accuracy gains against cost---choose bundles optimally subject to their budget constraints. Prior work [Chaudhury et al., 2026] studies such optimal pricing under the restriction that each dataset is assigned a linear price, and shows that computing optimal linear prices is computationally intractable. In contrast, we allow each dataset to be priced via a general function and show that this additional flexibility can not only increase the revenue but also restore tractability, yielding a surprising simultaneous improvement in economic performance and computational efficiency. Even when pricing functions are only required to be monotone and lower-continuous, optimal pricing admits a highly structured and simple form: each pricing function is piecewise linear and convex (PLC), and the optimal solution can be computed in polynomial time. Moreover, the total number of kinks across all pricing functions is bounded by the number of buyers. Consequently, when datasets significantly outnumber buyers, most pricing functions are effectively linear. We further empirically study the structure of optimal pricing by analyzing the number of kinks and the revenue gap between optimal nonlinear pricing and optimal linear pricing on simulations generated from a real dataset.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bhaskar Ray Chaudhury, Jugal Garg, Eklavya Sharma, Jiaxin Song. 2026-09-29. Non-Linear Pricing Restores Tractability for a Data Seller. https://arxiv.org/abs/2609.36589

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Honest Reporting in Scored Oversight: True-KL0 Property via the Prekopa Principle

We prove the True-KL$_0$ property for a parametric family of heterogeneous scoring rules arising in scored elicitation mechanisms (AI oversight, forecasting, expert surveys). An agent with private type $M>1$, scored through a $d$-dimensional outcome interface, reports to a principal who evaluates via a power-$p$ pseudospherical scoring rule, $p \in (d,d+1)$; $M$ captures the agent's information quality relative to a reference. Honest reporting is dominant-strategy optimal for every $d$ and every $p>1$, without a prior over the agent's type: a consequence of strict properness and identifiability, with a quadratic misreport-loss rate. True-KL$_0$, the property $R(M,p,d)<1$ for all $M>1$, $d \in \{2,3,4\}$, $p \in (d,d+1)$, is the quantitative core: $R$ is the Rayleigh quotient of the radial misreport channel of an annular oversight model, and True-KL$_0$ certifies a uniform curvature-domination margin for that channel: $1-R \ge 0.26$ ($R \le 0.7324$, semi-rigorous numerical certificate). Two structural tools drive the proof: (i) a substitution $y=(x+1)/(x-1)$ rewrites the loss integral $I_L$ as $\int_1^M F(y)(M^2-y^2)^{d/2} dy$ with $M$-independent weight $F(y)>0$; (ii) log-concavity of $I_L$ in $M$: algebraic for $d=2$ up to a small certified compact verification, via Prekopa's theorem plus semi-rigorous certificates for $d \in \{3,4\}$. True-KL$_0$ then follows from elementary tail bounds plus a certified bound on $M \in [1.001, 20]$. We also characterise the dimensional boundary: True-KL$_0$ holds for all $p \in (d,d+1)$ when $d \le 4$; $d=5$ is the unique transition, with $p_{crit}(5) \in [5.5718, 5.5750]$ (mpmath, not interval-certified); for $d=6,7$ (and conjecturally all $d \ge 6$) no threshold exists: the bound fails at every sampled $p \in (d,d+1)$.

cs.GT↗

Profit Reallocation Mechanisms in Tree-based Data Trading

Markets for data promise to unlock its economic value, yet in practice they remain far less active than expected---one reason is that those who supply data are not rewarded for the value it creates downstream. A defining feature of data is its \emph{replicability}: a buyer can refine purchased data into a new product and resell it to \emph{many} downstream buyers, so a single source seeds a branching cascade of resales that naturally forms a \emph{tree}. Because an upstream seller captures none of this downstream value, its incentive to trade is weakened. Existing works propose \emph{profit reallocation}---returning part of downstream revenue to upstream contributors---as a natural remedy. But whether profit reallocation works on the tree-structured markets that replicable data actually induces has remained open. To bridge this gap, we develop a principled framework for profit reallocation on tree-structured data markets. We introduce a sequential trading game on a tree and a general class of budget-feasible profit reallocation mechanisms (PRMs) over it. We derive efficient algorithms to compute the induced equilibria---a polynomial-time exact algorithm for discrete valuations and a fully polynomial-time approximation scheme (FPTAS) for continuous ones---via a subtree decomposition technique that tames the potential coupling across a seller's children. We then prove that, under mild assumptions, \emph{any} budget-feasible PRM weakly expands the trades that occur in equilibrium, and any budget-balanced PRM additionally weakly improves social welfare, relative to the baseline that reallocates nothing. Experiments on synthetic markets confirm that these benefits are substantial, persist even when the assumptions fail, and grow with the depth and branching of the tree.

cs.GT↗

From Reconnaissance to Response: Quantitative Risk Parameterization and Game Theoretic Containment in Modern Enterprise Attack

Modern Security Operations Centers struggle with delayed manual incident response, enabling adversaries to advance through the Cyber Kill Chain during early stage reconnaissance. While classical game theoretic defense models optimize strategic resource allocation, they rely on static utility matrices that fail to adapt to dynamic telemetry. This paper presents an integrated, metrics driven decision engine that bridges quantitative risk parameterization and continuous automated response time. Common Vulnerability Scoring Systems exploitability parameters are mapped to attacker success probabilities and evaluate defender log distributions via Factor Analysis of Information Risk Monte Carlo simulations. Real time SIEM logs streams are modeled as Poisson process arrival rates, dynamically updating defender posterior threat belief through sequential Bayesian filtering. A closed form threshold is derived by framing the interaction as a dynamic Bayesian Stackelberg game, where the expected unmitigated risk exceeds proactive containment cost. Parameterized against empirical data from the 2023 MGM Resorts and Caesars Entertainment cyber incident, simulation results demonstrate that the engine suppresses transient background noise while triggering automated SOAR network isolation within seconds of adversarial probing. Multi parameter sensitivity analysis confirms that the decision boundary dynamically adjusts to live perimeter vulnerability, offering a control theoretic foundation for sub minute automated threat containment.

cs.GT↗