Searcharxiv⌕ Search

arXiv · 2610.00760

LeanSide: A Formally Verified Co-Reasoning System for Natural-language Proofs

Abstract

Large language models are increasingly used as collaborators on deductive-reasoning tasks, but their outputs can hallucinate or pull users away from intended reasoning. Formal proof assistants provide machine-checked verification, but have a steep learning curve and require more granular reasoning than human written proofs. We explore an interface that combines these strengths, allowing users to write and revise free-form natural-language proofs while a verified backend checks their reasoning and returns feedback at the user's granularity. We study this interface in the context of undergraduate mathematics education by developing LeanSide, a formally verified co-reasoning system, which auto-formalizes student reasoning into Lean and informalizes verifier output into understandable feedback. We conducted user studies through classroom deployment and analyzed which system properties helped students make progress and which caused them to get stuck. We use these findings to derive design implications for using a formally verified backend in human-AI co-reasoning systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chenjun Guo, Manooshree Patel, Arnav Mehta, Krishiv Kothari, Thomas Lu, Niels Voss, Rayna Bhattacharyya, Peter Donovan, Bjoern Hartmann, Gireeja Ranade. 2026-09-30. LeanSide: A Formally Verified Co-Reasoning System for Natural-language Proofs. https://arxiv.org/abs/2610.00760

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AI Behavioral Science: A Framework and Agenda

We discuss the challenges and opportunities present in the rapidly emerging area of ``AI Behavioral Science.'' We frame it via three subfields. First, as AI becomes ubiquitous and is increasingly proprietary and opaque, it becomes vital to develop models of AI and methods for assessing AI behavior. We outline how tools developed to assess people's behaviors by social scientists can be used to model, assess and infer AI's behaviors biases, tendencies, and heuristics. Second, we also discuss how AI can change the ways in which we learn about human behavior. Beyond its computational power, AI offers new techniques for simulating, inferring, predicting, and analyzing human behaviors. Third, as humans and AI are interacting in increasingly complex and intertwined systems, we need to analyze and model human-AI interactions including how human and AI behaviors depend on interactions at the individual level, how interacting systems of humans and AI behave, and ultimately how AI's integration into society affects economic and political outcomes. We discuss current research, questions, agendas, and goals in each of these three subfields and how they depend upon each other.

cs.HC↗

Scaling Peer Assessments: An Integrity Report from a Large Engineering Internship

Assessing learning in large classrooms presents a significant challenge for individual instructors, who may have limited capacity to evaluate the understanding, participation, and assessment behaviour of every student. Peer assessments have been a way of distributing this responsibility among learners, allowing them to evaluate and provide feedback to one another while reducing dependence on instructor-led assessments. Building on this approach, we implemented a peer validation model within a large, multi-institutional internship programme in which students who demonstrated sufficient understanding were authorised to assess and validate their peers through short oral discussions. The assessment process began with the instructor validating a small group of students, who were then authorised to validate their peers, allowing the process to gradually expand across the cohort and operate at scale. This study examines how participants experienced the model and the extent to which assessment integrity was maintained, using an end-of-programme survey of 238 consenting respondents. Most participants regarded the activity as worthwhile, with 79.8% reporting that they solved problems they could not previously solve. However, 29.0% acknowledged at least one instance of reduced effort, a lowered validation standard, or reciprocal validation, while 88.7% believed that at least a little validation had occurred without proper examination. When asked how the process could be strengthened, participants selected post-validation discussion of solutions approximately twice as often as closer auditing or mentor-led validation. These findings provide descriptive evidence of both the potential and the integrity challenges of using peer validation as a scalable assessment approach in large learning environments.

cs.HC↗

Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI

Generative AI can support writing, but frictionless access may cause cognitive offloading before users develop their own ideas. We introduce Engage-to-Unlock, a productive-friction mechanism that unlocks generative capabilities after users meaningfully engage with the task. In a controlled experiment (N = 398), participants completed a writing task under one of four conditions: Human-Only, Standard Chatbot, Engage-to-Unlock, or Time-Matched Unlock, which matched unlock timing to Engage-to-Unlock participants but independent of users' engagement, then evaluated passages for evidence and inferential errors. Results show that Engage-to-Unlock redistributed effort across tasks: participants spent more time writing and less time evaluating, without increasing overall task duration. They also submitted more prompts than in other AI-assisted conditions and showed the highest accuracy-per-time evaluation efficiency across conditions. These findings suggest that designing GenAI access to encourage early human engagement may provide a productive form of friction, while retaining active AI use and efficient downstream evaluation.

cs.HC↗