TY - RPRT TI - Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks AU - Melissa Ailem AU - Katerina Marazopoulou AU - Charlotte Siska AU - James Bono PY - 2024 UR - https://arxiv.org/abs/2404.16966 ID - 2404.16966 ER -