arXiv · 2507.07029
Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices
Abstract
This paper presents the design and development of an OCR-powered pipeline for efficient table extraction from invoices. The system leverages Tesseract OCR for text recognition and custom post-processing logic to detect, align, and extract structured tabular data from scanned invoice documents. Our approach includes dynamic preprocessing, table boundary detection, and row-column mapping, optimized for noisy and non-standard invoice formats. The resulting pipeline significantly improves data extraction accuracy and consistency, supporting real-world use cases such as automated financial workflows and digital archiving.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Parshva Dhilankumar Patel. 2025-07-09. Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices. https://arxiv.org/abs/2507.07029
Cite the original work for its findings. Save a collection to share your selection of sources.