TY - RPRT TI - Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity AU - Zoë Prins AU - Samuele Punzo AU - Frank Wildenburg AU - Giovanni Cinà AU - Sandro Pezzelle PY - 2026 UR - https://arxiv.org/abs/2603.29396 ID - 2603.29396 ER -