SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
In Text-to-SQL tasks, large language models can generate structurally different SQL queries that return correct answers for the same intent. We call this phenomenon execution-correct structural divergence (ECSD). We introduce SQLStructEval, a framework that analyzes this behavior through canonical abstract syntax tree representations. We quantify structural diversity and agreement across generations. Experiments with different LLMs on multiple Text-to-SQL datasets, including Spider, document ECSD across models and datasets. Furthermore, our experiments demonstrate that generated queries are sensitive to question paraphrases and schema presentation. To address ECSD, we adopt a pipeline that first generates structured intermediate representations and then deterministically compiles them into SQL, improving execution accuracy and structural agreement among correct outputs. Structural analysis thus provides an additional diagnostic perspective that complements execution-based evaluation. Code is available at https://xanderzhou2022.github.io/AACL2026-SQLSTRUCTEVAL/.