TY - RPRT TI - Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard AU - Oguzhan Topsakal AU - Colby Jacob Edell AU - Jackson Bailey Harper PY - 2024 UR - https://arxiv.org/abs/2407.07796 ID - 2407.07796 ER -