arXiv · 1910.05535
From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction
Abstract
A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a great part of the automatic approaches have been relying on rules or supervised machine learning. We present a fully automatic unsupervised way of extracting parallel data for training a character-based sequence-to-sequence NMT (neural machine translation) model to conduct OCR error correction.
Explore related subjects
Keep this discovery
Mika Hämäläinen, Simon Hengchen. 2019-10-12. From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction. https://doi.org/10.26615/978-954-452-056-4_051
Cite the original work for its findings. Save a collection to share your selection of sources.