SearcharxivSearch

arXiv subjects

Mario Cortina-Borja

Publications and source records attributed to Mario Cortina-Borja.

6 recordsLinked to original sources

Designing Ambiguity-Aware Clerical Review: A Stratified Sampling Framework for Record Linkage and Deduplication

Clerical review of candidate record pairs remains the de facto gold standard for evaluating record linkage, but it is resource-intensive and often designed informally. We propose a design-based framework that treats clerical review as finite-population sampling over fine-grained strata defined by match weight, comparison pattern, record-level ambiguity, and demographic group. Match-probability bands are constructed from model-based score deciles. Within bands, strata combine comparison patterns with an ambiguity factor derived from matchability and conditional candidate perplexity. A band-specific margin-of-error profile encodes substantive priorities, such as tighter precision in high-score bands, while a single scaling parameter enforces the overall clerical budget. We evaluate the framework using a labelled dataset deduplicated in Splink, comprising 50,000 records and approximately 478,000 candidate pairs. We compare a baseline design reviewing about 23% of pairs with a budget-constrained design reviewing about 7%. The baseline accurately estimates global and band-specific match rates, while sampled distributions of comparison patterns, gender, and ambiguity broadly track the population. Under the budget design, global error approximately doubles, with the largest band-level errors in middle-score bands where matches, non-matches, and ambiguous cases are intermixed. Accuracy in the highest-score bands and the band-wise ambiguity profile are largely preserved, although representativeness by comparison pattern and gender declines. The framework generalises to other clerical-review objectives and can incorporate gold-standard data as prior information for design and calibration. It makes explicit the negotiable trade-offs between workload, precision, representativeness, and coverage of linkage uncertainty.

stat.ME

Cluster-based name embeddings reduce ethnic disparities in record linkage quality under realistic name corruption: evidence from the North Carolina Voter Registry

Differential ethnic-based record linkage errors can bias epidemiologic estimates. Prior evidence often conflates heterogeneity in error mechanisms with unequal exposure to error. Using snapshots of the North Carolina Voter Registry (Oct 2011-Oct 2022), we derived empirical name-discrepancy profiles to parameterise realistic corruptions. From an Oct 2022 extract (n=848,566), we generated five replicate corrupted datasets under three settings that separately varied mechanism heterogeneity and exposure inequality, and linked records back to originals using unadjusted Jaro-Winkler, Term Frequency (TF)-adjusted Jaro-Winkler, and a cluster-based forename-embedding comparator combined with TF-adjusted surname comparison. We evaluated false match rate (FMR), missed match rate (MMR) and white-centric disparities. At a fixed MMR near 0.20, overall error rates and ethnic disparities diverged substantially by model under disproportionate exposure to corruption. Term-frequency (TF)-adjusted Jaro-Winkler achieved very low overall FMR (0.55% (95% CI 0.54-0.57)) at overall MMR 20.34% (20.30-20.39), but large white-centric under-linkage disparities persisted: Hispanic voters had 36.3% (36.1-36.6) and Non-Hispanic Black voters 8.6% (8.6-8.7) higher FMRs compared to Non-Hispanic White groups. Relative to unadjusted string similarity, TF adjustment reduced these disparities (Hispanic: +60.4% (60.1-60.7) to +36.3%; Black: +13.1% (13.0-13.2) to +8.6%). The cluster-based forename-embedding model reduced missed-match disparities further (Hispanic: +10.2% (9.8-10.3); Black: +0.6% (0.4-0.7)), but at a cost of increasing overall FMR (4.28% (4.22-4.35)) at the same threshold. Unequal exposure to identifier error drove substantially larger disparities than mechanism heterogeneity alone; cluster-based embeddings markedly narrowed under-linkage disparities beyond TF adjustment.

stat.ME

Generalized network autoregressive modelling of longitudinal networks with application to presidential elections in the USA

Longitudinal networks are becoming increasingly relevant in the study of dynamic processes characterised by known or inferred community structure. Generalised Network Autoregressive (GNAR) models provide a parsimonious framework for exploiting the underlying network and multivariate time series. We introduce the community-$α$ GNAR model with interactions that exploits prior knowledge or exogenous variables for analysing interactions within and between communities, and can describe serial correlation in longitudinal networks. We derive new explicit finite-sample error bounds that validate analysing high-dimensional longitudinal network data with GNAR models, and provide insights into their attractive properties. We further illustrate our approach by analysing the dynamics of $\textit{Red, Blue}$ and $\textit{Swing}$ states throughout presidential elections in the USA from 1976 to 2020, that is, a time series of length twelve on 51 time series (US states and Washington DC). Our analysis connects network autocorrelation to eight-year long terms, highlights a possible change in the system after the 2016 election, and a difference in behaviour between $\textit{Red}$ and $\textit{Blue}$ states.

stat.ME

Decolonising Data Systems: Using Jyutping or Pinyin as tonal representations of Chinese names for data linkage

Data linkage is increasingly used in health research and policy making and is relied on for understanding health inequalities. However, linked data is only as useful as the underlying data quality, and differential linkage rates may induce selection bias in the linked data. A mechanism that selectively compromises data quality is name romanisation. Converting text of a different writing system into Latin based writing, or romanisation, has long been the standard process of representing names in character-based writing systems such as Chinese, Vietnamese, and other languages such as Swahili. Unstandardised romanisation of Chinese characters, due in part to problems of preserving the correct name orders the lack of proper phonetic representation of a tonal language, has resulted in poor linkage rates for Chinese immigrants. This opinion piece aims to suggests that the use of standardised romanisation systems for Cantonese (Jyutping) or Mandarin (Pinyin) Chinese, which incorporate tonal information, could improve linkage rates and accuracy for individuals with Chinese names. We used 771 Chinese and English names scraped from openly available sources, and compared the utility of Jyutping, Pinyin and the Hong Kong Government Romanisation system (HKG-romanisation) for representing Chinese names. We demonstrate that both Jyutping and Pinyin result in fewer errors compared with the HKG-romanisation system. We suggest that collecting and preserving people's names in their original writing systems is ethically and socially pertinent. This may inform development of language-specific pre-processing and linkage paradigms that result in more inclusive research data which better represents the targeted populations.

cs.CL

Modelling clusters in network time series with an application to presidential elections in the USA

Network time series are becoming increasingly relevant in the study of dynamic processes characterised by a known or inferred underlying network structure. Generalised Network Autoregressive (GNAR) models provide a parsimonious framework for exploiting the underlying network, even in the high-dimensional setting. We extend the GNAR framework by presenting the $\textit{community}$-$α$ GNAR model that exploits prior knowledge and/or exogenous variables for identifying and modelling dynamic interactions across communities in the network. We further analyse the dynamics of $\textit{ Red, Blue}$ and $\textit{Swing}$ states throughout presidential elections in the USA. Our analysis suggests interesting global and communal effects.

stat.ME

New tools for network time series with an application to COVID-19 hospitalisations

Network time series are becoming increasingly important across many areas in science and medicine and are often characterised by a known or inferred underlying network structure, which can be exploited to make sense of dynamic phenomena that are often high-dimensional. For example, the Generalised Network Autoregressive (GNAR) models exploit such structure parsimoniously. We use the GNAR framework to introduce two association measures: the network and partial network autocorrelation functions, and introduce Corbit (correlation-orbit) plots for visualisation. As with regular autocorrelation plots, Corbit plots permit interpretation of underlying correlation structures and, crucially, aid model selection more rapidly than using other tools such as AIC or BIC. We additionally interpret GNAR processes as generalised graphical models, which constrain the processes' autoregressive structure and exhibit interesting theoretical connections to graphical models via utilization of higher-order interactions. We demonstrate how incorporation of prior information is related to performing variable selection and shrinkage in the GNAR context. We illustrate the usefulness of the GNAR formulation, network autocorrelations and Corbit plots by modelling a COVID-19 network time series of the number of admissions to mechanical ventilation beds at 140 NHS Trusts in England & Wales. We introduce the Wagner plot that can analyse correlations over different time periods or with respect to external covariates. In addition, we introduce plots that quantify the relevance and influence of individual nodes. Our modelling provides insight on the underlying dynamics of the COVID-19 series, highlights two groups of geographically co-located `influential' NHS Trusts and demonstrates superior prediction abilities when compared to existing techniques.

stat.ME