arXiv · 2007.07082
Unsupervised Data Extraction from Computer-generated Documents with Single Line Formatting
Abstract
Processing large amounts of data is an essential problem of the big data era. Most of the data exchange is done via direct communication (using APIs) and well-structured file formats (JSON, XML, EDI, etc.), but a significant portion of the data is transferred using arbitrary formatted computer-generated documents (such as invoices, purchase orders, financial reports, etc.), which require sophisticated processing and human intervention for data interpretation and extraction. The currently available solutions, ranging from manual data entry to low-level scripting and data extraction tools, are costly and require human intervention. This paper describes the principle methodology for unsupervised, fully automatic data extraction from a wide range of computer-generated documents, assuming that their formatting reflects the original structure of the data sources. The presented methodology falls into the category of unsupervised machine learning and consists of the three main parts: (1) - detecting repeating patterns of text formatting by employing the relative feature space clustering and adaptive weighted feature score maps, (2) - detecting hierarchical formatting structures via collapsing and noise filtering procedure applied to the repeating formatting patterns and (3) - automatic configuration of the interactive data extraction tool (SiMX TextConverter) for fully automated processing.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vladimir Bernstein, Andrei Afanassenkov. 2020-07-15. Unsupervised Data Extraction from Computer-generated Documents with Single Line Formatting. https://arxiv.org/abs/2007.07082
Cite the original work for its findings. Save a collection to share your selection of sources.