TY - RPRT TI - Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AU - Xiaocan Li AU - Shiliang Wu AU - Zheng Shen PY - 2026 UR - https://arxiv.org/abs/2605.20402 ID - 2605.20402 ER -