SearcharxivSearch

arXiv subjects

Soheil Ghili

Publications and source records attributed to Soheil Ghili.

9 recordsLinked to original sources

Position: The Pre/Post-Training Boundary Should Govern IP in Industry-Academia ML Collaborations

Industry-academia ML collaborations routinely fail to launch -- not for scientific reasons, but because academics must publish while companies must protect models trained on proprietary data, and no standard contract framework resolves this tension. Because contracts are negotiated by legal departments alone, many apparent legal disputes are incentive misalignment problems that only scientists at the table can correctly diagnose. We propose PBOS (Protect-the-Business / Open-Source-the-Science), a community-adoptable contract template anchored to a single technically-grounded boundary: pre-training artifacts (architectures, training code, benchmarks, untrained weights) are open science; post-training artifacts (weights trained on proprietary data) are business IP. This boundary is technically meaningful, legally clean, and auditable -- and could not have been drawn correctly without scientists at the negotiating table. We argue the ML community should adopt PBOS as its default contract for such collaborations.

econ.GN

Training Language Models for Bilateral Trade with Private Information

Bilateral bargaining under incomplete information provides a controlled testbed for evaluating large language model (LLM) agent capabilities. Bilateral trade demands individual rationality, strategic surplus maximization, and cooperation to realize gains from trade. We develop a structured bargaining environment where LLMs negotiate via tool calls within an event-driven simulator, separating binding offers from natural-language messages to enable automated evaluation. The environment serves two purposes: as a benchmark for frontier models and as a training environment for open-weight models via reinforcement learning. In benchmark experiments, a round-robin tournament among five frontier models (15,000 negotiations) reveals that effective strategies implement price discrimination through sequential offers. Aggressive anchoring, calibrated concession, and temporal patience correlate with the highest surplus share and deal rate. Accommodating strategies that concede quickly disable price discrimination in the buyer role, yielding the lowest surplus capture and deal completion. Stronger models scale their behavior proportionally to item value, maintaining performance across price tiers; weaker models perform well only when wide zones of possible agreement offset suboptimal strategies. In training experiments, we fine-tune Qwen3 (8B, 14B) via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) against a fixed frontier opponent. These stages optimize competing objectives: SFT approximately doubles surplus share but reduces deal rates, while RL recovers deal rates but erodes surplus gains, reflecting the reward structure. SFT also compresses surplus variation across price tiers, which generalizes to unseen opponents, suggesting that behavioral cloning instills proportional strategies rather than memorized price points.

cs.GT

Pay-Per-Crawl Pricing for AI: The LM-Tree Agent

As AI systems shift from directing users to content toward consuming it directly, publishers need a new revenue model: charging AI crawlers for content access. This model, called pay-per-crawl, must solve a problem of mechanism selection at scale: content is too heterogeneous for a fixed pricing framework. Different sub-types warrant not only different price levels but different pricing rules based on different unstructured features, and there are too many to enumerate or design by hand. We propose the LM Tree, an adaptive pricing agent that grows a segmentation tree over the content library, using LLMs to discover what distinguishes high-value from low-value items and apply those attributes at scale, from binary purchase feedback alone. We evaluate the LM Tree on real content from a major German technology publisher, using 8,939 articles and 80,451 buyer queries with willingness-to-pay calibrated from actual AI crawler traffic. The LM Tree achieves a 65% revenue gain over a single static price and a 47% gain over two-category pricing, outperforming even the publisher's own 8-segment editorial taxonomy by 40% -- recovering content distinctions the publisher's own categories miss.

econ.GN

Auctions Meet Bandits: An Empirical Analysis

Sponsored search positions are typically allocated through real-time auctions, where the outcomes depend on advertisers' quality-adjusted bids - the product of their bids and quality scores. Although quality scoring helps promote ads with higher conversion outcomes, setting these scores for new advertisers in any given market is challenging, leading to the cold-start problem. To address this, platforms incorporate multi-armed bandit algorithms in auctions to balance exploration and exploitation. However, little is known about the optimal exploration strategies in such auction environments. We utilize data from a leading Asian mobile app store that places sponsored ads for keywords. The platform employs a Thompson Sampling algorithm within a second-price auction to learn quality scores and allocate a single sponsored position for each keyword. We empirically quantify the gains from optimizing exploration under this combined auction-bandit model and show that this problem differs substantially from the canonical bandit problem. Drawing on these empirical insights, we propose a customized exploration strategy in which the platform adjusts the exploration levels for each keyword according to its characteristics. We derive the Pareto frontier for revenue and efficiency and provide actionable policies, demonstrating substantial gains for the platform on both metrics when using a tailored exploration approach.

cs.GT

Second-degree Price Discrimination: Theoretical Analysis, Experiment Design, and Empirical Estimation

We build on theoretical results from the mechanism design literature to analyze empirical models of second-degree price discrimination (2PD). We show that for a random-coefficients discrete choice ("BLP") model to be suitable for studying 2PD, it must capture the covariance between two key random effects: (i) the "baseline" willingness to pay (affecting all product versions), and (ii) the perceived differentiation between versions. We then develop an experimental design that, among other features, identifies this covariance under common data constraints in 2PD environments. We implement this experiment in the field in collaboration with an international airline. Estimating the theoretically motivated empirical model on the experimental data, we demonstrate its applicability to 2PD decisions. We also show that test statistics from our design can enable qualitative inference on optimal 2PD policy even before estimating a demand model. Our methodology applies broadly across second-degree price discrimination settings.

econ.GN

An Empirical Analysis of Optimal Nonlinear Pricing in Business-to-Business Markets

In continuous-choice settings, consumers decide not only on whether to purchase a product, but also on how much to purchase. Thus, firms optimize a full price schedule rather than a single price point. This paper provides a methodology to empirically estimate the optimal schedule under multi-dimensional consumer heterogeneity with a focus on B2B applications. We apply our method to novel data from an educational-services firm that contains purchase-size information not only for deals that materialized, but also for potential deals that eventually failed. We show that this data, combined with identifying assumptions, helps infer how price sensitivity varies with "customer size". Using our estimated model, we show that the optimal second-degree price discrimination (i.e., optimal nonlinear tariff) improves the firm's profit upon linear pricing by at least 8.2%. That said, this second-degree price discrimination scheme only recovers 7.1% of the gap between the profitability of linear pricing and that of infeasible first degree price discrimination. We also conduct several further simulation analyses (i) empirically quantifying the magnitude by which incentive-compatibility constraints impact the optimal pricing and profits, (ii) comparing the role of demand- v.s. cost-side factors in shaping the optimal price schedule, and (iii) studying the implications of fixed fees for the optimal contract and profitability.

econ.GN

A Characterization for Optimal Bundling of Products with Non-Additive Values

This paper studies optimal bundling of products with non-additive values. Under monotonic preferences and single-peaked profits, I show a monopolist finds pure bundling optimal if and only if the optimal sales volume for the grand bundle is larger than the optimal sales volume for any smaller bundle. I then (i) detail how my analysis relates to "ratio monotonicity" results on bundling; and (ii) describe the implications for non-linear pricing.

econ.GN

Spatial Distribution of Supply and the Role of Market Thickness: Theory and Evidence from Ride Sharing

This paper studies the effects of economies of density in transportation markets, focusing on ridesharing. Our theoretical model predicts that (i) economies of density skew the supply of drivers away from less dense regions, (ii) the skew will be more pronounced for smaller platforms, and (iii) rideshare platforms do not find this skew efficient and thus use prices and wages to mitigate (but not eliminate) it. We then develop a general empirical strategy with simple implementation and limited data requirements to test for spatial skew of supply from demand. Applying our method to ride-level, multi-platform data from New York City (NYC), we indeed find evidence for a skew of supply toward busier areas, especially for smaller platforms. We discuss the implications of our analysis for business strategy (e.g., spatial pricing) and public policy (e.g., consequences of breaking up or downsizing a rideshare platform)

econ.GN

Eliminating Latent Discrimination: Train Then Mask

How can we control for latent discrimination in predictive models? How can we provably remove it? Such questions are at the heart of algorithmic fairness and its impacts on society. In this paper, we define a new operational fairness criteria, inspired by the well-understood notion of omitted variable-bias in statistics and econometrics. Our notion of fairness effectively controls for sensitive features and provides diagnostics for deviations from fair decision making. We then establish analytical and algorithmic results about the existence of a fair classifier in the context of supervised learning. Our results readily imply a simple, but rather counter-intuitive, strategy for eliminating latent discrimination. In order to prevent other features proxying for sensitive features, we need to include sensitive features in the training phase, but exclude them in the test/evaluation phase while controlling for their effects. We evaluate the performance of our algorithm on several real-world datasets and show how fairness for these datasets can be improved with a very small loss in accuracy.

cs.LG