arXiv · 2412.12362
How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games
Abstract
The deployment of large language models (LLMs) in diverse applications requires a thorough understanding of their decision-making strategies and behavioral patterns. As a supplement to a recent study on the behavioral Turing test, this paper presents a comprehensive analysis of five leading LLM-based chatbot families as they navigate a series of behavioral economics games. By benchmarking these AI chatbots, we aim to uncover and document both common and distinct behavioral patterns across a range of scenarios. The findings provide valuable insights into the strategic preferences of each LLM, highlighting potential implications for their deployment in critical decision-making roles.
Explore related subjects
Keep this discovery
Yutong Xie, Yiyao Liu, Zhuang Ma, Lin Shi, Xiyuan Wang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei. 2024-12-16. How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games. https://arxiv.org/abs/2412.12362
Cite the original work for its findings. Save a collection to share your selection of sources.