SearcharxivSearch

arXiv subjects

Weiran Shen

Publications and source records attributed to Weiran Shen.

At least 19 recordsLinked to original sources

Optimal Mediation Mechanism in Bilateral Trade

We study the problem of designing revenue-maximizing mechanisms for a selfish mediator who facilitates trade between a buyer and a seller. We consider a setting where the mediator does not have information advantage and the buyer's valuation is interdependent with the seller's private information. The mechanism may involve multi-round negotiations and flexible fee structures. We show that the mediator can restrict attention to a class of joint menu-selection mechanisms, where each mechanism can be represented as a two-dimensional menu. Each party privately selects an option from their own dimension and the two options together determine the menu entry. The mediator then recommends both parties whether to trade based on the jointly selected menu entry. We then establish an impossibility trilemma: no mechanism can simultaneously satisfy incentive compatibility, obedience, and informativeness. Motivated by this result, we characterize the optimal mechanisms under two relaxation conditions. First, when the seller's cost is constant, the optimal mechanism exhibits a threshold structure: trade occurs whenever the quality of the item exceeds a threshold that is decreasing in the buyer's type. Consequently, low-typed buyers receive more information, which in turn gives the mediator more power to charge from them. Second, when the mediator has veto power, the optimal mechanism also takes a threshold form, but in the opposite direction: trade occurs only if the quality falls below a threshold that is increasing in the buyer's type. As a result, items with lower qualities are more likely to be traded and the corresponding sellers benefit more, which discourages sellers of high qualities from participating and gives rise to a ``lemons market'' effect.

cs.GT

Inverse Contextual Bandits without Rewards: Learning from a Non-Stationary Learner via Suffix Imitation

We study the Inverse Contextual Bandit (ICB) problem, in which a learner seeks to optimize a policy while an observer, who cannot access the learner's rewards and only observes actions, aims to recover the underlying problem parameters. During the learning process, the learner's behavior naturally transitions from exploration to exploitation, resulting in non-stationary action data that poses significant challenges for the observer. To address this issue, we propose a simple and effective framework called Two-Phase Suffix Imitation. The framework discards data from an initial burn-in phase and performs empirical risk minimization using only data from a subsequent imitation phase. We derive a predictive decision loss bound that explicitly characterizes the bias-variance trade-off induced by the choice of burn-in length. Despite the severe information deficit, we show that a reward-free observer can achieve a convergence rate of $\tilde O(1/\sqrt{N})$, matching the asymptotic efficiency of a fully reward-aware learner. This result demonstrates that a passive observer can effectively uncover the optimal policy from actions alone, attaining performance comparable to that of the learner itself.

cs.LG

Enhancing Safety of Large Language Models via Embedding Space Separation

Large language models (LLMs) have achieved impressive capabilities, yet ensuring their safety against harmful prompts remains a critical challenge. Recent work has revealed that the latent representations (embeddings) of harmful and safe queries in LLMs typically exhibit linear separability, a property that has been exploited to construct attacks by perturbing the embeddings of harmful queries towards the safe subspace. Motivated by this observation, we propose a representation-level fine-tuning approach, named Embedding Space Separation (ES2), which improves LLM safety by explicitly enlarging the distance between harmful and safe representations in the embedding space. To prevent degradation of model's general capabilities, we introduce a Kullback-Leibler (KL) divergence regularization term into the loss function, which constrains the logits of the fine-tuned model to align with those of the original base model on harmless inputs. We evaluate our method on several open-source LLMs using standard safety benchmarks. Extensive experimental results demonstrate that our approach substantially improves model safety while maintaining comparable general capabilities.

cs.CL

Mechanism Design via Market Clearing-Prices for Value Maximizers under Budget and RoS Constraints

The transition to auto-bidding in online advertising has shifted the focus of auction theory from quasi-linear utility maximization to value maximization subject to financial constraints. We study mechanism design for buyers with private budgets and private Return-on-Spend (RoS) constraints, but public valuations, a setting motivated by modern advertising platforms where valuations are predicted via machine learning models. We introduce the extended Eisenberg-Gale program, a convex optimization framework generalized to incorporate RoS constraints. We demonstrate that the solution to this program is unique and characterizes the market's competitive equilibrium. Based on this theoretical analysis, we design a market-clearing mechanism and prove two key properties: (1) it is incentive-compatible with respect to financial constraints, making truthful reporting the optimal strategy; and (2) it achieves a tight 1/2-approximation of the first-best revenue benchmark, the maximum revenue of any feasible mechanism, regardless of IC. Finally, to enable practical implementation, we present a decentralized online algorithm. Ignoring logarithmic factors, we prove that under this algorithm, both the seller's revenue and each buyer's utility converge to the equilibrium benchmarks with a sublinear regret of $\tilde{O}(\sqrt{m})$ over $m$ auctions.

cs.GT

Enhancing Long-Term Welfare in Recommender Systems: An Information Revelation Approach

Improving the long-term user welfare (e.g., sustained user engagement) has become a central objective of recommender systems (RS). In real-world platforms, the creation behaviors of content creators plays a crucial role in shaping long-term welfare beyond short-term recommendation accuracy, making the effective steering of creator behavior essential to foster a healthier RS ecosystem. Existing works typically rely on re-ranking algorithms that heuristically adjust item exposure to steer creators' behavior. However, when embedded within recommendation pipelines, such a strategy often conflicts with the short-term objective of improving recommendation accuracy, leading to performance degradation and suboptimal long-term welfare. The well-established economics studies offer us valuable insights for an alternative approach without relying on recommendation algorithmic design: revealing information from an information-rich party (sender) to a less-informed party (receiver) can effectively change the receiver's beliefs and steer their behavior. Inspired by this idea, we propose an information-revealing framework, named Long-term Welfare Optimization via Information Revelation (LoRe). In this framework, we utilize a classical information revelation method (i.e., Bayesian persuasion) to map the stakeholders in RS, treating the platform as the sender and creators as the receivers. To address the challenge posed by the unrealistic assumption of traditional economic methods, we formulate the process of information revelation as a Markov Decision Process (MDP) and propose a learning algorithm trained and inferred in environments with boundedly rational creators. Extensive experiments on two real-world RS datasets demonstrate that our method can effectively outperform existing fair re-ranking methods and information revealing strategies in improving long-term user welfare.

cs.IR

Multiplayer General Lotto game

In this paper, we investigate the multiplayer General Lotto game across multiple battlefields, a significant variant of the Colonel Blotto game. In this version, each player employs a probability distribution for resource allocation, ensuring that their expected expenditure does not exceed their budget. We first establish the existence of the Nash equilibrium in a general setting, where players' budgets are asymmetric and the values of the battlefields are heterogeneous and asymmetric among players. Next, we provide a detailed characterization of the Nash equilibrium for multiple players on a single battlefield. In this characterization, we observe that the upper endpoints of the supports of players' equilibrium strategies coincide, and that the minimum value of a player's support above zero inversely correlates with his budget. We demonstrate the uniqueness of Nash equilibrium over a single battlefield in some scenarios. In the multi-battlefield setting, we prove that there is an upper bound on the average number of battlefields each player participates in. Additionally, we provide an example demonstrating the non-uniqueness of the Nash equilibrium in the context of multiple battlefields with multiple players. Finally, we present a solution for the Nash equilibrium in a symmetric case.

cs.GT

CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs

Game-playing ability serves as an indicator for evaluating the strategic reasoning capability of large language models (LLMs). While most existing studies rely on utility performance metrics, which are not robust enough due to variations in opponent behavior and game structure. To address this limitation, we propose \textbf{Cognitive Hierarchy Benchmark (CHBench)}, a novel evaluation framework inspired by the cognitive hierarchy models from behavioral economics. We hypothesize that agents have bounded rationality -- different agents behave at varying reasoning depths/levels. We evaluate LLMs' strategic reasoning through a three-phase systematic framework, utilizing behavioral data from six state-of-the-art LLMs across fifteen carefully selected normal-form games. Experiments show that LLMs exhibit consistent strategic reasoning levels across diverse opponents, confirming the framework's robustness and generalization capability. We also analyze the effects of two key mechanisms (Chat Mechanism and Memory Mechanism) on strategic reasoning performance. Results indicate that the Chat Mechanism significantly degrades strategic reasoning, whereas the Memory Mechanism enhances it. These insights position CHBench as a promising tool for evaluating LLM capabilities, with significant potential for future research and practical applications.

cs.AI

Stackelberg vs. Nash in the Lottery Colonel Blotto Game

Resource competition problems are often modeled using Colonel Blotto games, where players take simultaneous actions. However, many real-world scenarios involve sequential decision-making rather than simultaneous moves. To model these dynamics, we represent the Lottery Colonel Blotto game as a Stackelberg game, in which one player, the leader, commits to a strategy first, and the other player, the follower, responds. We derive the Stackelberg equilibrium for this game, formulating the leader's strategy as a bi-level optimization problem. To solve this, we develop a constructive method based on iterative game reductions, which allows us to efficiently compute the leader's optimal commitment strategy in polynomial time. Additionally, we identify the conditions under which the Stackelberg equilibrium coincides with the Nash equilibrium. Specifically, this occurs when the budget ratio between the leader and the follower equals a certain threshold, which we can calculate in closed form. In some instances, we observe that when the leader's budget exceeds this threshold, both players achieve higher utilities in the Stackelberg equilibrium compared to the Nash equilibrium. Lastly, we show that, in the best case, the leader can achieve an infinite utility improvement by making an optimal first move compared to the Nash equilibrium.

cs.GT

IBCB: Efficient Inverse Batched Contextual Bandit for Behavioral Evolution History

Traditional imitation learning focuses on modeling the behavioral mechanisms of experts, which requires a large amount of interaction history generated by some fixed expert. However, in many streaming applications, such as streaming recommender systems, online decision-makers typically engage in online learning during the decision-making process, meaning that the interaction history generated by online decision-makers includes their behavioral evolution from novice expert to experienced expert. This poses a new challenge for existing imitation learning approaches that can only utilize data from experienced experts. To address this issue, this paper proposes an inverse batched contextual bandit (IBCB) framework that can efficiently perform estimations of environment reward parameters and learned policy based on the expert's behavioral evolution history. Specifically, IBCB formulates the inverse problem into a simple quadratic programming problem by utilizing the behavioral evolution history of the batched contextual bandit with inaccessible rewards. We demonstrate that IBCB is a unified framework for both deterministic and randomized bandit policies. The experimental results indicate that IBCB outperforms several existing imitation learning algorithms on synthetic and real-world data and significantly reduces running time. Additionally, empirical analyses reveal that IBCB exhibits better out-of-distribution generalization and is highly effective in learning the bandit policy from the interaction history of novice experts.

cs.LG

Optimal Fixed-Price Mechanism with Signaling

Consider a trade market with one seller and multiple buyers. The seller aims to sell an indivisible item and maximize their revenue. This paper focuses on a simple and popular mechanism--the fixed-price mechanism. Unlike the standard setting, we assume there is information asymmetry between buyers and the seller. Specifically, we allow the seller to design information before setting the fixed price, which implies that we study the mechanism design problem in a broader space. We call this mechanism space the fixed-price signaling mechanism. We assume that buyers' valuation of the item depends on the quality of the item. The seller can privately observe the item's quality, whereas buyers only see its distribution. In this case, the seller can influence buyers' valuations by strategically disclosing information about the item's quality, thereby adjusting the fixed price. We consider two types of buyers with different levels of rationality: ex-post individual rational (IR) and ex-interim individual rational. We show that when the market has only one buyer, the optimal revenue generated by the fixed-price signaling mechanism is identical to that of the fixed-price mechanism, regardless of the level of rationality. Furthermore, when there are multiple buyers in the market and all of them are ex-post IR, we show that there is no fixed-price mechanism that is obedient for all buyers. However, if all buyers are ex-interim IR, we show that the optimal fixed-price signaling mechanism will generate more revenue for the seller than the fixed-price mechanism.

cs.GT

Selling an Item through Persuasion

A monopolistic seller aims to sell an indivisible item to multiple potential buyers. Each buyer's valuation depends on their private type and the item's quality. The seller can observe the quality but it is unknown to buyers. This quality information is valuable to buyers, so it is beneficial for the seller to strategically design experiments that reveal information about the quality before deciding to sell the item to whom and at what price. We study the problem of designing a revenue-maximizing mechanism that allows the seller to disclose information and sell the item. First, we recast the revelation principle to our setting, showing that the seller can focus on one-round mechanisms without loss of generality. We then formulate the mechanism design problem as an optimization problem and derive the optimal solution in closed form. The optimal mechanism includes a set of experiments and payment functions. After eliciting buyers' types, the optimal mechanism asks a buyer to buy and sets a price accordingly. The optimal information structure involves partitioning the quality space. Additionally, we show that our results can be extended to a broader class of distributions and valuation functions.

cs.GT

Revenue Maximization Mechanisms for an Uninformed Mediator with Communication Abilities

Consider a market where a seller owns an item for sale and a buyer wants to purchase it. Each player has private information, known as their type. It can be costly and difficult for the players to reach an agreement through direct communication. However, with a mediator as a trusted third party, both players can communicate privately with the mediator without worrying about leaking too much or too little information. The mediator can design and commit to a multi-round communication protocol for both players, in which they update their beliefs about the other player's type. The mediator cannot force the players to trade but can influence their behaviors by sending messages to them. We study the problem of designing revenue-maximizing mechanisms for the mediator. We show that the mediator can, without loss of generality, focus on a set of direct and incentive-compatible mechanisms. We then formulate this problem as a mathematical program and provide an optimal solution in closed form under a regularity condition. Our mechanism is simple and has a threshold structure. Additionally, we extend our results to general cases by utilizing a variant version of the ironing technique. In the end, we discuss some interesting properties revealed from the optimal mechanism, such as, in the optimal mechanism, the mediator may even lose money in some cases.

cs.GT

SRAP-Agent: Simulating and Optimizing Scarce Resource Allocation Policy with LLM-based Agent

Public scarce resource allocation plays a crucial role in economics as it directly influences the efficiency and equity in society. Traditional studies including theoretical model-based, empirical study-based and simulation-based methods encounter limitations due to the idealized assumption of complete information and individual rationality, as well as constraints posed by limited available data. In this work, we propose an innovative framework, SRAP-Agent (Simulating and Optimizing Scarce Resource Allocation Policy with LLM-based Agent), which integrates Large Language Models (LLMs) into economic simulations, aiming to bridge the gap between theoretical models and real-world dynamics. Using public housing allocation scenarios as a case study, we conduct extensive policy simulation experiments to verify the feasibility and effectiveness of the SRAP-Agent and employ the Policy Optimization Algorithm with certain optimization objectives. The source code can be found in https://github.com/jijiarui-cather/SRAPAgent_Framework

cs.CL

LTP-MMF: Towards Long-term Provider Max-min Fairness Under Recommendation Feedback Loops

Multi-stakeholder recommender systems involve various roles, such as users, and providers. Previous work pointed out that max-min fairness (MMF) is a better metric to support weak providers. However, when considering MMF, the features or parameters of these roles vary over time, how to ensure long-term provider MMF has become a significant challenge. We observed that recommendation feedback loops (named RFL) will greatly influence the provider MMF in the long term. RFL means that recommender systems can only receive feedback on exposed items from users and update recommender models incrementally based on this feedback. When utilizing the feedback, the recommender model will regard the unexposed items as negative. In this way, the tail provider will not get the opportunity to be exposed, and its items will always be considered negative samples. Such phenomena will become more and more serious in RFL. To alleviate the problem, this paper proposes an online ranking model named Long-Term Provider Max-min Fairness (named LTP-MMF). Theoretical analysis shows that the long-term regret of LTP-MMF enjoys a sub-linear bound. Experimental results on three public recommendation benchmarks demonstrated that LTP-MMF can outperform the baselines in the long term.

cs.IR

AAVDiff: Experimental Validation of Enhanced Viability and Diversity in Recombinant Adeno-Associated Virus (AAV) Capsids through Diffusion Generation

Recombinant adeno-associated virus (rAAV) vectors have revolutionized gene therapy, but their broad tropism and suboptimal transduction efficiency limit their clinical applications. To overcome these limitations, researchers have focused on designing and screening capsid libraries to identify improved vectors. However, the large sequence space and limited resources present challenges in identifying viable capsid variants. In this study, we propose an end-to-end diffusion model to generate capsid sequences with enhanced viability. Using publicly available AAV2 data, we generated 38,000 diverse AAV2 viral protein (VP) sequences, and evaluated 8,000 for viral selection. The results attested the superiority of our model compared to traditional methods. Additionally, in the absence of AAV9 capsid data, apart from one wild-type sequence, we used the same model to directly generate a number of viable sequences with up to 9 mutations. we transferred the remaining 30,000 samples to the AAV9 domain. Furthermore, we conducted mutagenesis on AAV9 VP hypervariable regions VI and V, contributing to the continuous improvement of the AAV9 VP sequence. This research represents a significant advancement in the design and functional validation of rAAV vectors, offering innovative solutions to enhance specificity and transduction efficiency in gene therapy applications.

cs.AI

Deep Generative Modeling on Limited Data with Regularization by Nontransferable Pre-trained Models

Deep generative models (DGMs) are data-eager because learning a complex model on limited data suffers from a large variance and easily overfits. Inspired by the classical perspective of the bias-variance tradeoff, we propose regularized deep generative model (Reg-DGM), which leverages a nontransferable pre-trained model to reduce the variance of generative modeling with limited data. Formally, Reg-DGM optimizes a weighted sum of a certain divergence and the expectation of an energy function, where the divergence is between the data and the model distributions, and the energy function is defined by the pre-trained model w.r.t. the model distribution. We analyze a simple yet representative Gaussian-fitting case to demonstrate how the weighting hyperparameter trades off the bias and the variance. Theoretically, we characterize the existence and the uniqueness of the global minimum of Reg-DGM in a non-parametric setting and prove its convergence with neural networks trained by gradient-based methods. Empirically, with various pre-trained feature extractors and a data-dependent energy function, Reg-DGM consistently improves the generation performance of strong DGMs with limited data and achieves competitive results to the state-of-the-art methods. Our implementation is available at https://github.com/ML-GSAI/Reg-ADA-APA.

cs.LG

Sequential Persuasion Using Limited Experiments

Bayesian persuasion and its derived information design problem has been one of the main research agendas in the economics and computation literature over the past decade. However, when attempting to apply its model and theory, one is often limited by the fact that the sender can only implement very restricted information structures. Moreover, in this case, the sender can possibly achieve higher expected utility by performing a sequence of feasible experiments, where the choice of each experiment depends on the outcomes of all previous experiments. Indeed, it has been well observed that real life persuasions often take place in rounds during which the sender exhibits experiments/arguments sequentially. We study the sender's expected utility maximization using finite and infinite sequences of experiments. For infinite sequences of experiments, we characterize the supremum of the sender's expected utility using a function that generalizes the concave closure definition in the standard Bayesian persuasion problem. With this characterization, we first study a special case where the sender can use feasible experiments to achieve the optimal expected utility of the standard Bayesian persuasion without feasibility constraints, which is a trivial utility upper bound, and establish structural findings about the sender's optimal sequential design in this case. Then we derive conditions under which the sender's optimal sequential design exists; when an optimal sequential design exists, there exists an optimal design that is Markovian, i.e., the choice of the next experiment only depends on the receiver's current belief.

cs.GT

P-MMF: Provider Max-min Fairness Re-ranking in Recommender System

In this paper, we address the issue of recommending fairly from the aspect of providers, which has become increasingly essential in multistakeholder recommender systems. Existing studies on provider fairness usually focused on designing proportion fairness (PF) metrics that first consider systematic fairness. However, sociological researches show that to make the market more stable, max-min fairness (MMF) is a better metric. The main reason is that MMF aims to improve the utility of the worst ones preferentially, guiding the system to support the providers in weak market positions. When applying MMF to recommender systems, how to balance user preferences and provider fairness in an online recommendation scenario is still a challenging problem. In this paper, we proposed an online re-ranking model named Provider Max-min Fairness Re-ranking (P-MMF) to tackle the problem. Specifically, P-MMF formulates provider fair recommendation as a resource allocation problem, where the exposure slots are considered the resources to be allocated to providers and the max-min fairness is used as the regularizer during the process. We show that the problem can be further represented as a regularized online optimizing problem and solved efficiently in its dual space. During the online re-ranking phase, a momentum gradient descent method is designed to conduct the dynamic re-ranking. Theoretical analysis showed that the regret of P-MMF can be bounded. Experimental results on four public recommender datasets demonstrated that P-MMF can outperformed the state-of-the-art baselines. Experimental results also show that P-MMF can retain small computationally costs on a corpus with the large number of items.

cs.IR