$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $π^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from the collected tables together with relevant metadata, generating complex reasoning questions whose answers are automatically determined and validated through dual-path code execution, 3) finally, back-translating chain-of-thoughts solutions grounded in realistic context. Supervised fine-tuning with gpt-oss-20b and Qwen3-4B-Instruct-2507 on $π^2$ yields consistent improvements across four long-context reasoning benchmarks and our alike $π^2$-Bench, with average absolute accuracy gains of +6.25% and +3.37% respectively. Through deeper analyses, we observe that reasoning style contributes little, while faithful reasoning patterns discovered by back translation and grounded realistic long context, as $π^2$ is designed for, are crucial for the improvement. Our code, data, and models are fully open-source at https://github.com/vtpss/pi-squared.