arXiv · 2604.12105
Automated BPMN Model Generation from Textual Process Descriptions: A Multi-Stage LLM-Driven Approach
Abstract
Automatically reconstructing BPMN models from unstructured natural-language descriptions remains challenging due to heterogeneous modeling conventions, multilingual sources, and the lack of reliable ground truth. We present a scalable, multi-stage LLM-driven pipeline that automates both ground-truth construction and model reconstruction. Multilingual BPMN XML files are translated into English, validated using execution-oriented compliance checks in SpiffWorkflow, and iteratively repaired through targeted LLM-guided corrections to produce a consistent ground-truth corpus. From these validated models, process descriptions are generated and used to reconstruct executable BPMN~2.0 XML diagrams without manual curation. We introduce a multi-dimensional similarity framework combining structural metrics, type-distribution alignment, and embedding-based semantic measures. In an empirical study of 750 public BPMN diagrams, the pipeline generated 387 validated ground-truth models and achieved average reconstruction similarity above 0.75, including approximately 50 near-perfect reconstructions differing only in minor naming variations. The results demonstrate that LLMs can generate structurally compliant and semantically meaningful BPMN diagrams at scale.
Explore related subjects
Keep this discovery
Ion Matei, Maksym Zhenirovskyy, Praveen Kumar Menaka Sekar, Hon Yung Wong. 2026-04-13. Automated BPMN Model Generation from Textual Process Descriptions: A Multi-Stage LLM-Driven Approach. https://arxiv.org/abs/2604.12105
Cite the original work for its findings. Save a collection to share your selection of sources.