SearcharxivSearch

arXiv subjects

Antoine Isaac

Publications and source records attributed to Antoine Isaac.

8 recordsLinked to original sources

Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections

Cultural Heritage (CH) data hold invaluable knowledge, reflecting the history, traditions, and identities of societies, and shaping our understanding of the past and present. However, many CH collections contain outdated or offensive descriptions that reflect historical biases. CH Institutions (CHIs) face significant challenges in curating these data due to the vast scale and complexity of the task. To address this, we develop an AI-powered tool that detects offensive terms in CH metadata and provides contextual insights into their historical background and contemporary perception. We leverage a multilingual vocabulary co-created with marginalized communities, researchers, and CH professionals, along with traditional NLP techniques and Large Language Models (LLMs). Available as a standalone web app and integrated with major CH platforms, the tool has processed over 7.9 million records, contextualizing the contentious terms detected in their metadata. Rather than erasing these terms, our approach seeks to inform, making biases visible and providing actionable insights for creating more inclusive and accessible CH collections.

cs.CL

Knowledge Graphs in the Libraries and Digital Humanities Domain

Knowledge graphs represent concepts (e.g., people, places, events) and their semantic relationships. As a data structure, they underpin a digital information system, support users in resource discovery and retrieval, and are useful for navigation and visualization purposes. Within the libaries and humanities domain, knowledge graphs are typically rooted in knowledge organization systems, which have a century-old tradition and have undergone their digital transformation with the advent of the Web and Linked Data. Being exposed to the Web, metadata and concept definitions are now forming an interconnected and decentralized global knowledge network that can be curated and enriched by community-driven editorial processes. In the future, knowledge graphs could be vehicles for formalizing and connecting findings and insights derived from the analysis of possibly large-scale corpora in the libraries and digital humanities domain.

cs.DL

Recommendations for the Technical Infrastructure for Standardized International Rights Statements

This white paper is the product of a joint Digital Public Library of America (DPLA)-Europeana working group organized to develop minimum rights statement metadata standards for organizations that contribute to DPLA and Europeana. This white paper deals specifically with the technical infrastructure of a common namespace (rightsstatements.org) that hosts the rights statements to be used by (at minimum) the DPLA and Europeana. These recommendations for a common technical infrastructure for rights statements outline a simple, flexible, and extensible framework to host the rights statements at rightsstatements.org. This white paper specifically outlines the management of rights statements as linked open data. The rights statements are published according to Best Practices for Publishing RDF Vocabularies. They are encoded into dereferenceable URIs, express further information encoded in RDF, and link to existing vocabularies and standards. The rights statements adhere to expressions of existing rights vocabularies. Furthermore the paper reviews the publication and implementation to make the rights statements available through human-readable web pages augmented with machine-readable formats.

cs.DL

Rightsstatements.org White Paper: Requirements for the Technical Infrastructure for Standardized International Rights Statements

This document is part of the deliverables created by the RightsStatements.org consortium. It provides the technical requirements for implementation of the Standardized International Rights Statements. These requirements are based on the principles and specifications found in the normative Recommendations for Standardized International Rights Statements. This document replaces and supersedes the previously released Recommendations for the Technical Infrastructure for Standardized Rights Statements, released by this working group. The Requirements for the Technical Infrastructure for Standardized International Rights Statements describes the expected behaviours for a service that enables the delivery of human and machine-readable representations of the rights statements. It documents the fundamental decisions that informed the development of a data model grounded in Linked Data approaches. This document also provides proposed implementation guidelines and a non-normative set of examples for incorporating rights statements into provider metadata.

cs.DL

Achieving interoperability between the CARARE schema for monuments and sites and the Europeana Data Model

Mapping between different data models in a data aggregation context always presents significant interoperability challenges. In this paper, we describe the challenges faced and solutions developed when mapping the CARARE schema designed for archaeological and architectural monuments and sites to the Europeana Data Model (EDM), a model based on Linked Data principles, for the purpose of integrating more than two million metadata records from national monument collections and databases across Europe into the Europeana digital library.

cs.DL

Hierarchical structuring of Cultural Heritage objects within large aggregations

Huge amounts of cultural content have been digitised and are available through digital libraries and aggregators like Europeana.eu. However, it is not easy for a user to have an overall picture of what is available nor to find related objects. We propose a method for hier- archically structuring cultural objects at different similarity levels. We describe a fast, scalable clustering algorithm with an automated field selection method for finding semantic clusters. We report a qualitative evaluation on the cluster categories based on records from the UK and a quantitative one on the results from the complete Europeana dataset.

cs.DL

Finding Quality Issues in SKOS Vocabularies

The Simple Knowledge Organization System (SKOS) is a standard model for controlled vocabularies on the Web. However, SKOS vocabularies often differ in terms of quality, which reduces their applicability across system boundaries. Here we investigate how we can support taxonomists in improving SKOS vocabularies by pointing out quality issues that go beyond the integrity constraints defined in the SKOS specification. We identified potential quantifiable quality issues and formalized them into computable quality checking functions that can find affected resources in a given SKOS vocabulary. We implemented these functions in the qSKOS quality assessment tool, analyzed 15 existing vocabularies, and found possible quality issues in all of them.

cs.DL

LCSH, SKOS and Linked Data

A technique for converting Library of Congress Subject Headings MARCXML to Simple Knowledge Organization System (SKOS) RDF is described. Strengths of the SKOS vocabulary are highlighted, as well as possible points for extension, and the integration of other semantic web vocabularies such as Dublin Core. An application for making the vocabulary available as linked-data on the Web is also described.

cs.DL