SearcharxivSearch

arXiv subjects

Renfei Tan

Publications and source records attributed to Renfei Tan.

3 recordsLinked to original sources

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect on behavioral diversity in sequential decision-making remains under-explored. We study this question in a controlled suite of deterministic board games based on tic-tac-toe variants, where optimal actions are exactly computable and diversity can be measured directly. Across state-level evaluation, arena gameplay, and training trajectories, we find that reasoning-mode generation frequently suppresses action diversity without uniformly improving action accuracy. Furthermore, standard SFT improves accuracy but often induces premature diversity collapse, which exceeds what is minimally required by the accuracy-diversity tradeoff. We then show that action augmentation, which trains on all optimal actions per state rather than a single demonstrated action, would partially mitigates this effect. Our results identify narrow-support imitation as a source of policy collapse in LLM decision-making and suggest that preserving action support during SFT is important for maintaining exploratory behavior.

cs.CL

Multi-agent Adaptive Mechanism Design

We study a sequential mechanism design problem in which a principal seeks to elicit truthful reports from multiple rational agents while starting with no prior knowledge of agents' beliefs. We introduce Distributionally Robust Adaptive Mechanism (DRAM), a general framework combining insights from both mechanism design and online learning to jointly address truthfulness and cost-optimality. Throughout the sequential game, the mechanism estimates agents' beliefs and iteratively updates a distributionally robust linear program with shrinking ambiguity sets to reduce payments while preserving truthfulness. Our mechanism guarantees truthful reporting with high probability while achieving $\tilde{O}(\sqrt{T})$ cumulative regret, and we establish a matching lower bound showing that no feasible adaptive mechanism can asymptotically do better. The framework generalizes to plug-in estimators, supporting structured priors and delayed feedback. To our knowledge, this is the first adaptive mechanism under general settings that maintains truthfulness and achieves optimal regret when incentive constraints are unknown and must be learned.

cs.GT

Payoff Control in Repeated Games

Evolutionary game theory is a powerful mathematical framework to study how intelligent individuals adjust their strategies in collective interactions. It has been widely believed that it is impossible to unilaterally control players' payoffs in games, since payoffs are jointly determined by all players. Until recently, a class of so-called zero-determinant strategies are revealed, which enables a player to make a unilateral payoff control over her partners in two-action repeated games with a constant continuation probability. The existing methods, however, lead to the curse of dimensionality when the complexity of games increases. In this paper, we propose a new mathematical framework to study ruling strategies (with which a player unilaterally makes a linear relation rule on players' payoffs) in repeated games with an arbitrary number of actions or players, and arbitrary continuation probability. We establish an existence theorem of ruling strategies and develop an algorithm to find them. In particular, we prove that strict Markov ruling strategy exists only if either the repeated game proceeds for an infinite number of rounds, or every round is repeated with the same probability. The proposed mathematical framework also enables the search of collaborative ruling strategies for an alliance to control outsiders. Our method provides novel theoretical insights into payoff control in complex repeated games, which overcomes the curse of dimensionality.

math.OC