Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
Lean verifies that a generated declaration is well typed, but not that it states what the user intended. For statement autoformalization without canonical Lean targets, we study two questions: how far an LLM-based semantic criterion can be trusted, and how much compilation overstates faithfulness across systems. Our criterion requires Lean compilation and agreement of GPT-5.2 and Gemini-2.5-Pro. On an independently audited random sample, it agrees with the human majority on 91.5\% of cases (Wilson 95\% CI: 81.6--96.3\%); humans confirm 95.7\% of the outputs it accepts and 77--81\% of the outputs it rejects, so the criterion is reliable in aggregate and errs on the conservative side. A comparison with LeanScorer, an independent third-family judge, and a BEq formal cross-check support the same picture. Across eight systems evaluated on 227 graduate-level statements, the compile--faithfulness gap varies widely with the system, from 1.3 percentage points for one-shot Gemini-2.5-Pro to 29.5 points for a tool-augmented GPT-5.2 agent, which compiles 87.2\% of statements but is faithful on 57.7\%. A $2^3$ tool ablation of this agent shows that Lean feedback drives most of its gain in compilation, and nearly half of that gain consists of outputs that fail the semantic criterion.