SearcharxivSearch

arXiv subjects

U. Springmann

Publications and source records attributed to U. Springmann.

3 recordsLinked to original sources

OCR of historical printings with an application to building diachronic corpora: A case study using the RIDGES herbal corpus

This article describes the results of a case study that applies Neural Network-based Optical Character Recognition (OCR) to scanned images of books printed between 1487 and 1870 by training the OCR engine OCRopus [@breuel2013high] on the RIDGES herbal text corpus [@OdebrechtEtAlSubmitted]. Training specific OCR models was possible because the necessary *ground truth* is available as error-corrected diplomatic transcriptions. The OCR results have been evaluated for accuracy against the ground truth of unseen test sets. Character and word accuracies (percentage of correctly recognized items) for the resulting machine-readable texts of individual documents range from 94% to more than 99% (character level) and from 76% to 97% (word level). This includes the earliest printed books, which were thought to be inaccessible by OCR methods until recently. Furthermore, OCR models trained on one part of the corpus consisting of books with different printing dates and different typesets *(mixed models)* have been tested for their predictive power on the books from the other part containing yet other fonts, mostly yielding character accuracies well above 90%. It therefore seems possible to construct generalized models trained on a range of fonts that can be applied to a wide variety of historical printings still giving good results. A moderate postcorrection effort of some pages will then enable the training of individual models with even better accuracies. Using this method, diachronic corpora including early printings can be constructed much faster and cheaper than by manual transcription. The OCR methods reported here open up the possibility of transforming our printed textual cultural heritage into electronic text by largely automatic means, which is a prerequisite for the mass conversion of scanned books.

cs.CL

Automatic quality evaluation and (semi-) automatic improvement of OCR models for historical printings

Good OCR results for historical printings rely on the availability of recognition models trained on diplomatic transcriptions as ground truth, which is both a scarce resource and time-consuming to generate. Instead of having to train a separate model for each historical typeface, we propose a strategy to start from models trained on a combined set of available transcriptions in a variety of fonts. These \emph{mixed models} result in character accuracy rates over 90\% on a test set of printings from the same period of time, but without any representation in the training data, demonstrating the possibility to overcome the typography barrier by generalizing from a few typefaces to a larger set of (similar) fonts in use over a period of time. The output of these mixed models is then used as a baseline to be further improved by both fully automatic methods and semi-automatic methods involving a minimal amount of manual transcriptions. In order to evaluate the recognition quality of each model in a series of models generated during the training process in the absence of any ground truth, we introduce two readily observable quantities that correlate well with true accuracy. These quantities are \emph{mean character confidence C} (as given by the OCR engine OCRopus) and \emph{mean token lexicality L} (a distance measure of OCR tokens from modern wordforms taking historical spelling patterns into account, which can be calculated for any OCR engine). Whereas the fully automatic method is able to improve upon the result of a mixed model by only 1-2 percentage points, already 100-200 hand-corrected lines lead to much better OCR results with character error rates of only a few percent. This procedure minimizes the amount of ground truth production and does not depend on the previous construction of a specific typographic model.

cs.DL

Atmospheric NLTE-Models for the Spectroscopic Analysis of Blue Stars with Winds. II. Line-Blanketed Models

We present new or improved methods for calculating NLTE, line-blanketed model atmospheres for hot stars with winds (spectral types A to O), with particular emphasis on a fast performance. These methods have been implemented into a previous, more simple version of the model atmosphere code FASTWIND (Santolaya-Rey et al.1997) and allow to spectroscopically analyze rather large samples of massive stars in a reasonable time-scale, using state-of-the-art physics. We describe our (partly approximate) approach to solve the equations of statistical equilibrium for those elements which are primarily responsible for line-blocking and blanketing, as well as an approximate treatment of the line-blocking itself, which is based on a simple statistical approach using suitable means for line opacities and emissivities. Furthermore, we comment on our implementation of a consistent temperature structure. In the second part, we concentrate on a detailed comparison with results from those two codes which have been used in alternative spectroscopical investigations, namely CMFGEN (Hillier & Miller 1998) and WM-Basic (Pauldrach et al. 2001). All three codes predict almost identical temperature structures and fluxes for lambda > 400 A, whereas at lower wavelengths a number of discrepancies are found. Optical H/He lines as synthesized by FASTWIND are compared with results from CMFGEN, obtaining a remarkable coincidence, except for the HeI singlets in the temperature range between 36,000 to 41,000 K for dwarfs and between 31,000 to 35,000 K for supergiants, where CMFGEN predicts much weaker lines. Consequences due to these discrepancies are discussed.

astro-ph