Searcharxiv⌕ Search

arXiv subjects

Torben S. D. Johansen

Publications and source records attributed to Torben S. D. Johansen.

5 recordsLinked to original sources

How to deal with machine learning bias in economic history

Machine learning (ML) has rapidly transformed economic history, lowering costs of digitization, data linkage, and imputation, and making information in historical text usable at scale. This paper offers a practical guide to using these tools well. However, ML tools have also created new problems. Prediction errors are often systematically correlated with covariates of interest, so even highly accurate models can distort and sometimes reverse coefficients, and standard validation cannot detect this. Given that ML tools often perform worse for historical data, this problem is especially severe for the field of economic history. We also identify a solution to this problem. We show that recent debiasing methods can correct such bias for a wide class of applications, using a small, randomly sampled set of expert-coded labels while retaining the efficiency of large-scale prediction. We organize the field with a taxonomy of three ML tasks, survey the literature along it, and indicate where debiasing applies and where validation against proxies remains the only recourse. We close with best-practice guidance on digitization, model choice, and reproducibility.

econ.GN↗

Group-Level Treatment Effect Heterogeneity in Difference-in-Differences: A Balanced Approach

Understanding how treatment effects vary across groups is central to policy evaluation. In Difference-in-Differences designs, heterogeneity is often studied using subgroup or triple-difference analyses, which can suffer from conservative inference, reliance on parametric interaction structures, and sensitivity to differences in covariate distributions across groups. We propose the Balanced Group Average Treatment Effect on the Treated (BGATT), a new estimand that isolates heterogeneity in treatment responses from differences in covariate composition and is identified under standard conditional parallel-trends assumptions. BGATT provides a transparent target for comparing group-specific treatment effects. We derive an influence-function representation and develop estimators that are $\sqrt{n}$-consistent and asymptotically normal under flexible machine-learning estimation of high-dimensional nuisance components, enabling valid inference on both group-specific effects and differences across groups. Simulation evidence shows favorable finite-sample performance.

econ.EM↗

Optimal Treatment Allocation under Constraints

In optimal policy problems where treatment effects vary at the individual level, optimally allocating treatments to recipients is complex even when potential outcomes are known. We present an algorithm for multi-arm treatment allocation problems that is guaranteed to find the optimal allocation in strongly polynomial time, and which is able to handle arbitrary potential outcomes as well as constraints on treatment requirement and capacity. Further, starting from an arbitrary allocation, we show how to optimally re-allocate treatments in a Pareto-improving manner. To showcase our results, we use data from Danish nurse home visiting for infants. We estimate nurse specific treatment effects for children born 1959-1967 in Copenhagen, comparing nurses against each other. We exploit random assignment of newborn children to nurses within a district to obtain causal estimates of nurse-specific treatment effects using causal machine learning. Using these estimates, and treating the Danish nurse home visiting program as a case of an optimal treatment allocation problem (where a treatment is a nurse), we document room for significant productivity improvements by optimally re-allocating nurses to children. Our estimates suggest that optimal allocation of nurses to children could have improved average yearly earnings by USD 1,815 and length of education by around two months.

econ.EM↗

DARE: A large-scale handwritten date recognition system

Handwritten text recognition for historical documents is an important task but it remains difficult due to a lack of sufficient training data in combination with a large variability of writing styles and degradation of historical documents. While recurrent neural network architectures are commonly used for handwritten text recognition, they are often computationally expensive to train and the benefit of recurrence drastically differs by task. For these reasons, it is important to consider non-recurrent architectures. In the context of handwritten date recognition, we propose an architecture based on the EfficientNetV2 class of models that is fast to train, robust to parameter choices, and accurately transcribes handwritten dates from a number of sources. For training, we introduce a database containing almost 10 million tokens, originating from more than 2.2 million handwritten dates which are segmented from different historical documents. As dates are some of the most common information on historical documents, and with historical archives containing millions of such documents, the efficient and automatic transcription of dates has the potential to lead to significant cost-savings over manual transcription. We show that training on handwritten text with high variability in writing styles result in robust models for general handwritten text recognition and that transfer learning from the DARE system increases transcription accuracy substantially, allowing one to obtain high accuracy even when using a relatively small training sample.

cs.CV↗

Applications of Machine Learning in Document Digitisation

Data acquisition forms the primary step in all empirical research. The availability of data directly impacts the quality and extent of conclusions and insights. In particular, larger and more detailed datasets provide convincing answers even to complex research questions. The main problem is that 'large and detailed' usually implies 'costly and difficult', especially when the data medium is paper and books. Human operators and manual transcription have been the traditional approach for collecting historical data. We instead advocate the use of modern machine learning techniques to automate the digitisation process. We give an overview of the potential for applying machine digitisation for data collection through two illustrative applications. The first demonstrates that unsupervised layout classification applied to raw scans of nurse journals can be used to construct a treatment indicator. Moreover, it allows an assessment of assignment compliance. The second application uses attention-based neural networks for handwritten text recognition in order to transcribe age and birth and death dates from a large collection of Danish death certificates. We describe each step in the digitisation pipeline and provide implementation insights.

cs.CV↗