SearcharxivSearch

arXiv subjects

Jonathan Bourne

Publications and source records attributed to Jonathan Bourne.

9 recordsLinked to original sources

The Character Error Vector: Decomposable errors for page-level OCR evaluation

The Character Error Rate (CER) is a key metric for evaluating the quality of Optical Character Recognition (OCR). However, this metric assumes that text has been perfectly parsed, which is often not the case. Under page-parsing errors, CER becomes undefined, limiting its use as a metric and making evaluating page-level OCR challenging, particularly when using data that do not share a labelling schema. We introduce the Character Error Vector (CEV), a bag-of-characters evaluator for OCR. The CEV can be decomposed into parsing and OCR, and interaction error components. This decomposability allows practitioners to focus on the part of the Document Understanding pipeline that will have the greatest impact on overall text extraction quality. The CEV can be implemented using a variety of methods, of which we demonstrate SpACER (Spatially Aware Character Error Rate) and a Character distribution method using the Jensen-Shannon Distance. We validate the CEV's performance against other metrics: first, the relationship with CER; then, parse quality; and finally, as a direct measure of page-level OCR quality. The validation process shows that the CEV is a valuable bridge between parsing metrics and local metrics like CER. We analyse a dataset of archival newspapers made of degraded images with complex layouts and find that state-of-the-art end-to-end models are outperformed by more traditional pipeline approaches. Whilst the CEV requires character-level positioning for optimal triage, thresholding on easily available values can predict the main error source with an F1 of 0.91. We provide the CEV as part of a Python library to support Document understanding research.

cs.CV

The COTe score: A decomposable framework for evaluating Document Layout Analysis models

Document Layout analysis (DLA), is the process by which a page is parsed into meaningful elements, often using machine learning models. Typically, the quality of a model is judged using general object detection metrics such as IoU, F1 or mAP. However, these metrics are designed for images that are 2D projections of 3D space, not for the natively 2D imagery of printed media. This discrepancy can result in misleading or uninformative interpretation of model performance by the metrics. To encourage more robust, comparable, and nuanced DLA, we introduce: The Structural Semantic Unit (SSU) a relational labelling approach that shifts the focus from the physical to the semantic structure of the content; and the Coverage, Overlap, Trespass, and Excess (COTe) score, a decomposable metric for measuring page parsing quality. We demonstrate the value of these methods through case studies and by evaluating 5 common DLA models on 3 DLA datasets. We show that the COTe score is more informative than traditional metrics and reveals distinct failure modes across models, such as breaching semantic boundaries or repeatedly parsing the same region. In addition, the COTe score reduces the interpretation-performance gap by up to 76% relative to the F1. Notably, we find that the COTe's granularity robustness largely holds even without explicit SSU labelling, lowering the barriers to entry for using the system. Finally, we release an SSU labelled dataset and a Python library for applying COTe in DLA projects.

cs.CV

Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models

Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism of Oscar Wilde's 19th century often has no or poor quality Optical Character Recognition (OCR), reducing the accessibility of these archives and making them unreadable both figuratively and literally. This paper helps address the issue by performing OCR on "The Nineteenth Century Serials Edition" (NCSE), an 84k-page collection of 19th-century English newspapers and periodicals, using Pixtral 12B, a pre-trained image-to-text language model. The OCR capability of Pixtral was compared to 4 other OCR approaches, achieving a median character error rate of 1%, 5x lower than the next best model. The resulting NCSE v2.0 dataset features improved article identification, high-quality OCR, and text classified into four types and seventeen topics. The dataset contains 1.4 million entries, and 321 million words. Example use cases demonstrate analysis of topic similarity, readability, and event tracking. NCSE v2.0 is freely available to encourage historical and sociological research. As a result, 21st-century readers can now share Oscar Wilde's disappointment with 19th-century journalistic standards, reading the unreadable from the comfort of their own computers.

cs.CL

Scrambled text: training Language Models to correct OCR errors using synthetic data

OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these errors using the context provided by the corrupted text and the broader socio-cultural context, a process called Context Leveraging OCR Correction (CLOCR-C). However, getting sufficient training data for fine-tuning such models can prove challenging. This paper shows that fine-tuning a language model on synthetic data using an LM and using a character level Markov corruption process can significantly improve the ability to correct OCR errors. Models trained on synthetic data reduce the character error rate by 55% and word error rate by 32% over the base LM and outperform models trained on real data. Key findings include; training on under-corrupted data is better than over-corrupted data; non-uniform character level corruption is better than uniform corruption; More tokens-per-observation outperforms more observations for a fixed token budget. The outputs for this paper are a set of 8 heuristics for training effective CLOCR-C models, a dataset of 11,000 synthetic 19th century newspaper articles and scrambledtext a python library for creating synthetic corrupted data.

cs.CL

CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models

The digitisation of historical print media archives is crucial for increasing accessibility to contemporary records. However, the process of Optical Character Recognition (OCR) used to convert physical records to digital text is prone to errors, particularly in the case of newspapers and periodicals due to their complex layouts. This paper introduces Context Leveraging OCR Correction (CLOCR-C), which utilises the infilling and context-adaptive abilities of transformer-based language models (LMs) to improve OCR quality. The study aims to determine if LMs can perform post-OCR correction, improve downstream NLP tasks, and the value of providing the socio-cultural context as part of the correction process. Experiments were conducted using seven LMs on three datasets: the 19th Century Serials Edition (NCSE) and two datasets from the Overproof collection. The results demonstrate that some LMs can significantly reduce error rates, with the top-performing model achieving over a 60\% reduction in character error rate on the NCSE dataset. The OCR improvements extend to downstream tasks, such as Named Entity Recognition, with increased Cosine Named Entity Similarity. Furthermore, the study shows that providing socio-cultural context in the prompts improves performance, while misleading prompts lower performance. In addition to the findings, this study releases a dataset of 91 transcribed articles from the NCSE, containing a total of 40 thousand words, to support further research in this area. The findings suggest that CLOCR-C is a promising approach for enhancing the quality of existing digital archives by leveraging the socio-cultural information embedded in the LMs and the text requiring correction.

cs.CL

What's in the laundromat? Mapping and characterising offshore owned domestic property in London

The UK, particularly London, is a global hub for money laundering, a significant portion of which uses domestic property. However, understanding the distribution and characteristics of offshore domestic property in the UK is challenging due to data availability. This paper attempts to remedy that situation by enhancing a publicly available dataset of UK property owned by offshore companies. We create a data processing pipeline which draws on several datasets and machine learning techniques to create a parsed set of addresses classified into six use classes. The enhanced dataset contains 138,000 properties 44,000 more than the original dataset. The majority are domestic (95k), with a disproportionate amount of those in London (42k). The average offshore domestic property in London is worth 1.33 million GBP collectively this amounts to approximately 56 Billion GBP. We perform an in-depth analysis of the offshore domestic property in London, comparing the price, distribution and entropy/concentration with Airbnb property, low-use/empty property and conventional domestic property. We estimate that the total amount of offshore, low-use and airbnb property in London is between 144,000 and 164,000 and that they are collectively worth between 145-174 billion GBP. Furthermore, offshore domestic property is more expensive and has higher entropy/concentration than all other property types. In addition, we identify two different types of offshore property, nested and individual, which have different price and distribution characteristics. Finally, we release the enhanced offshore property dataset, the complete low-use London dataset and the pipeline for creating the enhanced dataset to reduce the barriers to studying this topic.

cs.LG

High Tension Lines: Predicting robustness of high-voltage power-grids to cascading failure using network embedding

This paper explores whether graph embedding methods can be used as a tool for analysing the robustness of power-grids within the framework of network science. The paper focuses on the strain elevation tension spring embedding (SETSe) algorithm and compares it to node2vec and Deep Graph Infomax, and the measures mean edge capacity and line load. These five methods are tested on how well they can predict the collapse point of the giant component of a network under random attack. The analysis uses seven power-grid networks, ranging from 14 to 2000 nodes. In total, 3456 load profiles are created for each network by loading the edges of the network to have a range of tolerances and concentrating network capacity into fewer edges. One hundred random attack sequences are generated for each load profile, and the mean number of attacks required for the giant component to collapse for each profile is recorded. The relationship between the embedding values for each load profile and the mean collapse point is then compared across all five methods. It is found that only SETSe and line load perform well as proxies for robustness with $R^2 = 0.89$ for both measures. When tested on a time series normal operating conditions line load performs exceptionally well ($R=0.99$). However, the SETSe algorithm provides valuable qualitative insight into the state of the power-grid by leveraging its method local smoothing and global weighting of node features to provide an interpretable geographical embedding. This paper shows that graph representation algorithms can be used to analyse network properties such as robustness to cascading failure attacks, even when the network is embedded at node level.

eess.SY

The spring bounces back: Introducing the Strain Elevation Tension Spring embedding algorithm for network representation

This paper introduces the Strain Elevation Tension Spring embedding (SETSe) algorithm, a graph embedding method that uses a physics model to create node and edge embeddings in undirected attribute networks. Using a low-dimensional representation, SETSe is able to differentiate between graphs that are designed to appear identical using standard network metrics such as number of nodes, number of edges and assortativity. The embeddings generated position the nodes such that sub-classes, hidden during the embedding process, are linearly separable, due to the way they connect to the rest of the network. SETSe outperforms five other common graph embedding methods on both graph differentiation and sub-class identification. The technique is applied to social network data, showing its advantages over assortativity as well as SETSe's ability to quantify network structure and predict node type. The algorithm has a convergence complexity of around $\mathcal{O}(n^2)$, and the iteration speed is linear ($\mathcal{O}(n)$), as is memory complexity. Overall, SETSe is a fast, flexible framework for a variety of network and graph tasks, providing analytical insight and simple visualisation for complex systems.

cs.SI

Don't go chasing artificial waterfalls: Simulating cascading failures in the power grid and the impact of artificial line-limit methods on results

Research into cascading failures in power-transmission networks requires detailed data on the capacity of individual transmission lines. However, these data are often unavailable to researchers. As a result, line limits are often modelled by assuming they are proportional to some average load. Little research exists, however, to support this assumption as being realistic. In this paper, we analyse the proportional-loading (PL) approach and compare it to two linear models that use voltage and initial power flow as variables. In conducting this modelling, we test the ability of artificial line limits to model true line limits, the damage done during an attack and the order in which edges are lost. we also test how accurately these methods rank the relative performance of different attack strategies. We find that the linear models are the top-performing method or close to the top in all tests. In comparison, the tolerance value that produces the best PL limits changes depending on the test. The PL approach was a particularly poor fit when the line tolerance was less than two, which is the most commonly used value range in cascading-failure research. We also find indications that the accuracy of modelling line limits does not indicate how well a model will represent grid collapse. In addition, we find evidence that the network's topology can be used to estimate the system's true mean loading. The findings of this paper provide an understanding of the weaknesses of the PL approach and offer an alternative method of line-limit modelling.

eess.SY