arXiv · 2510.18112
Does Reasoning Help LLM Agents Play Dungeons and Dragons? A Prompt Engineering Experiment
Abstract
This paper explores the application of Large Language Models (LLMs) and reasoning to predict Dungeons & Dragons (DnD) player actions and format them as Avrae Discord bot commands. Using the FIREBALL dataset, we evaluated a reasoning model, DeepSeek-R1-Distill-LLaMA-8B, and an instruct model, LLaMA-3.1-8B-Instruct, for command generation. Our findings highlight the importance of providing specific instructions to models, that even single sentence changes in prompts can greatly affect the output of models, and that instruct models are sufficient for this task compared to reasoning models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Patricia Delafuente, Arya Honraopatil, Lara J. Martin. 2025-10-20. Does Reasoning Help LLM Agents Play Dungeons and Dragons? A Prompt Engineering Experiment. https://arxiv.org/abs/2510.18112
Cite the original work for its findings. Save a collection to share your selection of sources.