Searcharxiv⌕ Search

arXiv subjects

Andrea Scharnhorst

Publications and source records attributed to Andrea Scharnhorst.

At least 19 recordsLinked to original sources

Knowledge Tectonics: A Geodynamic Inspired Framework for Modeling Epistemic Changes

Emerging knowledge can be envisioned as protruding magma. This paper describes how such a metaphoric view can be turned into a model to capture the 'continental drift' of concepts in an epistemic lithosphere. We call this new approach Knowledge Tectonics. We detail conceptual, mathematical and engineering operations to create such a scalable framework which allows us to interpret and manage knowledge evolution within Semantic Web environments. We use the WikiArt Emotions dataset which contains information on 4,105 paintings spanning 600 years to construct a proof--of--concept interface which enables visual analytics of where artworks are situated in a specific landscape of features. We demonstrate how fused semantic and pragmatic metadata can be modelled as evolving pressure zones. This way we are making 'forces' behind the evolution of artifacts of creativity visible. Core elements of our new workflow are gradient vector fields derived from Poisson potential surfaces applied onto dynamic knowledge graphs, eventually capturing stylistic and emotional shifts as directed intensity flows. By treating artefacts of creative processes as a dynamic manifold, we provide a novel methodology for quantifying the 'drift' of human inquiry. We argue that this approach is applicable also to other areas of creative human actions, including scientific knowledge production.

cs.DL↗

Co-creation of AI technology, empowering curators of cultural heritage information and guarding research commons

The substance of this paper is the description of the use of Retrieval-Augmented Generation (RAG) for specific digital collections of cultural assets. The collections are provided by institutions operating in the cultural sector. The topical areas are the humanities and social sciences. More concretely, most of the work presented here was enabled by a European-funded research project MuseIT which is clearly situated in the realm of fostering new technologies for Cultural Heritage. We adhere to this interaction by presenting a sequence of our experimentations. This sequence is narrated as a specific journey of engineering all executed around a specific data-sharing and archiving platform Dataverse. Implementing a local chatbot for collections - a method also known as RAG in Information Retrieval - is the current culmination of this journey. The engineering journey we describe in the core of the paper starts from "archives for everyone" and ends with "local chatbots for specific collections".

cs.DL↗

Chatting with Papers: A Hybrid Approach Using LLMs and Knowledge Graphs

This demo paper reports on a new workflow \textit{GhostWriter} that combines the use of Large Language Models and Knowledge Graphs (semantic artifacts) to support navigation through collections. Situated in the research area of Retrieval Augmented Generation, this specific workflow represents the creation of local and adaptable chatbots. Based on the tool-suite \textit{EverythingData} at the backend, \textit{GhostWriter} provides an interface that enables querying and ``chatting'' with a collection. Applied iteratively, the workflow supports the information needs of researchers when interacting with a collection of papers, whether it be to gain an overview, to learn more about a specific concept and its context, and helps the researcher ultimately to refine their research question in a controlled way. We demonstrate the workflow for a collection of articles from the \textit{method data analysis} journal published by GESIS -- Leibniz-Institute for the Social Sciences. We also point to further application areas.

cs.DL↗

A Knowledge Base for Arts and Inclusion -- The Dataverse data archival platform as a knowledge base management system enabling multimodal accessibility

Creating an inclusive art environment requires engaging multiple senses for a fully immersive experience. Culture is inherently synesthetic, enriched by all senses within a shared time and space. In an optimal synesthetic setting, people of all abilities can connect meaningfully; when one sense is compromised, other channels can be enhanced to compensate. This is the power of multimodality. Digital technology is increasingly able to capture aspects of multimodality. To document multimodality aspects of cultural practices and products for the long-term remains a challenge. Many artistic products from the performing arts tend to be multimodal, and are often immersive, so only a multimodal repository can offer a platform for this work. To our knowledge there is no single, comprehensive repository with a knowledge base to serve arts and disability. By knowledge base, we mean classifications, taxonomies, or ontologies (in short, knowledge organisation systems). This paper presents innovative ways to develop a knowledge base which capture multimodal features of archived representations of cultural assets, but also indicate various forms how to interact with them including machine-readable description. We will demonstrate how back-end and front-end applications, in a combined effort, can support accessible archiving and data management for complex digital objects born out of artistic practices and make them available for wider audiences.

cs.DL↗

Measuring a moving target -- Innovation studies in practice

This paper pays a tribute to Loet's work in a specific way. More than 20 years ago Loet Leydesdorff and myself designed a programme for future innovation studies 'Measuring the knowledge base - a programme of innovation studies'. Although, the funding programme we envisioned eventually did not materialise, the proposal text set out the main lines of our research collaboration over the coming decades. This paper revisits main statements of this programme and discusses their remaining validity in the light of more recent research. Core to the Leydesdorff and Scharnhorst text was a system-theoretical, evolutionary perspective on science dynamics, newly emerging structures and phenomena and addressed the question to which extent they could be meaningful studied using quantitative approaches. This paper looks into three cases - all examples of newly emerging institutional structures and related practices in science. They are all located at the interface between research and research infrastructures. While discussing how the programmatic ideas written up at the beginning of the 2000s still informs measurement attempts in those three cases, the paper also touches upon questions of epistemological foundations of quantitative studies in general. The main conclusion is that combining measurement experiments with philosophical reflection remains important.

cs.DL↗

Fostering Data Communities -- perspective from a Data Archive Service Provider

This paper aims to bridge between the current scientific discourse about the dynamics of data communities in research infrastructures and practical experiences at a data archive which provides services for such data communities. We describe and analyse policies and practices within DANS-KNAW, the Dutch national centre of expertise and repository for research data concerning the interaction with communities in general. We take the case of the emerging DANS Data Station Life Sciences to study how a data archive navigates between observation of data research needs and anticipation of research data archival solutions. This paper offers a unique view of the complex dynamics between data communities (including lay experts) and data service providers. It adds nuances to understanding the emergence of a data community and the role of data service providers, both supporting and shaping, in this process.

cs.DL↗

Publishing a Knowledge Organization System as Linked Data: The Case of the Universal Decimal Classification

Linked data (LD) technology is hailed as a long-awaited solution in web-based information exchange. Linked Open Data (LOD) bring this to another level by enabling meaningful linking of resources and creating a global, openly accessible knowledge graph. Our case is the Universal Decimal Classification (UDC) and the challenges for a KOS service provider to maintain an LD service. UDC was created during the period 1896--1904 to support systematic organization and information retrieval of a bibliography. When discussing UDC as LD we make a distinction between two types of UDC data or two provenances: UDC source data, and UDC codes as they appear in metadata. To serve the purpose of supplying semantics one has to front--end UDC LD with a service that can parse and interpret complex UDC strings. While the use of UDC is free the publishing and distributing of UDC data is protected by a licence. Publishing of UDC both as LD and as LOD must be provided for within a complex service that would allow open access as well as access through a paywall barrier for different levels of licences. The practical task of publishing the UDC as LOD was informed by the '10Things guidelines'. The process includes conceptual parts and technological parts. The transition to a new technology is never a purely mechanical act but is a research endeavour in its own right. The UDC case has shown the importance of cross-domain, inter-disciplinary collaboration which needs experts well situated in multiple knowledge domains.

cs.DL↗

The Need for Knowledge Organization. Introduction to the book Linking Knowledge: Linked Open Data for Knowledge Organization

This book is not restricted to semantic web (SW) technologies. An aspiration was to contribute to the awakening of a dialogue between information and documentation concerned with knowledge organization systems (KOSs), and branches in computer science with an emphasis on machines, algorithms and ontologies. The technological evolution of the last decades has not only fostered the emergence of ever more KOSs but also semantic web technologies. Both the actions of 'making a KOS' and 'applying existing KOSs' represent research. The design of an information layer for a knowledge domain and the design of a domain specific research process are intrinsically interwoven. We extended our intervention to KOS practices into education, by presenting a translation of existing standards and recommendations about linked open data (LOD) publishing for non-experts. The chapters describe the state of the art in providing KOSs as semantic artefacts; how the state of the art is applied in new fields; how the state of the art is pushed towards new technological solutions by being confronted with new applications; how best practices need to be tailored towards specific solutions; and what challenges occur when merging new and old ways of expressing KOSs. The linked data (LD) ecosystem represents a source of knowledge generation, acquisition, production and dissemination. The underlying discourse shows historical vision alongside the promise of linking knowledge for interaction. The already maturing ecosystems of the SW are interlocking information institutions clearly devoted to the expansion of human experience through the growth of knowledge interaction.

cs.DL↗

Classifications as Linked Open Data. Challenges and Opportunities

Linked Data (LD) as a web--based technology enables in principle the seamless, machine--supported integration, interplay and augmentation of all kinds of knowledge, into what has been labeled a huge knowledge graph. Despite decades of web technology and, more recently, the LD approach, the task to fully exploit these new technologies in the public domain is only commencing. One specific challenge is to transfer techniques developed preweb to order our knowledge into the realm of Linked Open Data (LOD) This paper illustrates two different models in which a general analytico--synthetic classification can be published and made available as LD. In both cases, an LD solution deals with the intricacies of a pre--coordinated indexing language.

cs.DL↗

Lost or found? Discovering data needed for research

Finding data is a necessary precursor to being able to reuse data, although relatively little large-scale empirical evidence exists about how researchers discover, make sense of and (re)use data for research. This study presents evidence from the largest known survey investigating how researchers discover and use data that they do not create themselves. We examine the data needs and discovery strategies of respondents, propose a typology for data reuse and probe the role of social interactions and literature search in data discovery. We consider how data communities can be conceptualized according to data uses and propose practical applications of our findings for designers of data discovery systems and repositories. Specifically, we consider how to design for a diversity of practices, how communities of use can serve as an entry point for design and the role of metadata in supporting both sensemaking and social interactions.

cs.DL↗

Searching Data: A Review of Observational Data Retrieval Practices in Selected Disciplines

A cross-disciplinary examination of the user behaviours involved in seeking and evaluating data is surprisingly absent from the research data discussion. This review explores the data retrieval literature to identify commonalities in how users search for and evaluate observational research data. Two analytical frameworks rooted in information retrieval and science technology studies are used to identify key similarities in practices as a first step toward developing a model describing data retrieval.

cs.DL↗

Understanding Data Search as a Socio-technical Practice

Open research data are heralded as having the potential to increase effectiveness, productivity, and reproducibility in science, but little is known about the actual practices involved in data search. The socio-technical problem of locating data for reuse is often reduced to the technological dimension of designing data search systems. We combine a bibliometric study of the current academic discourse around data search with interviews with data seekers. In this article, we explore how adopting a contextual, socio-technical perspective can help to understand user practices and behavior and ultimately help to improve the design of data discovery systems.

cs.DL↗

Digital Data Archives as Knowledge Infrastructures: Mediating Data Sharing and Reuse

Digital archives are the preferred means for open access to research data. They play essential roles in knowledge infrastructures - robust networks of people, artifacts, and institutions - but little is known about how they mediate information exchange between stakeholders. We open the "black box" of data archives by studying DANS, the Data Archiving and Networked Services institute of The Netherlands, which manages 50+ years of data from the social sciences, humanities, and other domains. Our interviews, weblogs, ethnography, and document analyses reveal that a few large contributors provide a steady flow of content, but most are academic researchers who submit datasets infrequently and often restrict access to their files. Consumers are a diverse group that overlaps minimally with contributors. Archivists devote about half their time to aiding contributors with curation processes and half to assisting consumers. Given the diversity and infrequency of usage, human assistance in curation and search remains essential. DANS' knowledge infrastructure encompasses public and private stakeholders who contribute, consume, harvest, and serve their data - many of whom did not exist at the time the DANS collections originated - reinforcing the need for continuous investment in digital data archives as their communities, technologies, and services evolve.

cs.DL↗

Exploration of reproducibility issues in scientometric research Part 1: Direct reproducibility

This is the first part of a small-scale explorative study in an effort to start assessing reproducibility issues specific to scientometrics research. This effort is motivated by the desire to generate empirical data to inform debates about reproducibility in scientometrics. Rather than attempt to reproduce studies, we explore how we might assess "in principle" reproducibility based on a critical review of the content of published papers. The first part of the study focuses on direct reproducibility - that is the ability to reproduce the specific evidence produced by an original study using the same data, methods, and procedures. The second part (Velden et al. 2018) is dedicated to conceptual reproducibility - that is the robustness of knowledge claims towards verification by an alternative approach using different data, methods and procedures. The study is exploratory: it investigates only a very limited number of publications and serves us to develop instruments for identifying potential reproducibility issues of published studies: These are a categorization of study types and a taxonomy of threats to reproducibility. We work with a select sample of five publications in scientometrics covering a variation of study types of theoretical, methodological, and empirical nature. Based on observations made during our exploratory review, we conclude this paper with open questions on how to approach and assess the status of direct reproducibility in scientometrics, intended for discussion at the special track on "Reproducibility in Scientometrics" at STI2018 in Leiden.

cs.DL↗

Exploration of Reproducibility Issues in Scientometric Research Part 2: Conceptual Reproducibility

This is the second part of a small-scale explorative study in an effort to assess reproducibility issues specific to scientometrics research. This effort is motivated by the desire to generate empirical data to inform debates about reproducibility in scientometrics. Rather than attempt to reproduce studies, we explore how we might assess "in principle" reproducibility based on a critical review of the content of published papers. While the first part of the study (Waltman et al. 2018) focuses on direct reproducibility - that is the ability to reproduce the specific evidence produced by an original study using the same data, methods, and procedures, this second part is dedicated to conceptual reproducibility - that is the robustness of knowledge claims towards verification by an alternative approach using different data, methods and procedures. The study is exploratory: it investigates only a very limited number of publications and serves us to develop instruments for identifying potential reproducibility issues of published studies: These are a categorization of study types and a taxonomy of threats to reproducibility. We work with a select sample of five publications in scientometrics covering a variation of study types of theoretical, methodological, and empirical nature. Based on observations made during our exploratory review, we conclude with open questions on how to approach and assess the status of conceptual reproducibility in scientometrics intended for discussion at the special track on "Reproducibility in Scientometrics" at STI2018 in Leiden.

cs.DL↗

Connecting KOSs and the LOD Cloud

This paper describes a specific project, the current situation leading to it, its project design and first results. In particular, we will examine the terminology employed in the Linked Open Data cloud and compare this to the terminology employed in both the Universal Decimal Classification and the Basic Concepts Classification. We will explore whether these classifications can encourage greater consistency in LOD terminology. We thus hope to link the largely distinct scholarly literatures that address LOD and KOSs.

cs.DL↗

Contextualization of topics: Browsing through the universe of bibliographic information

This paper describes how semantic indexing can help to generate a contextual overview of topics and visually compare clusters of articles. The method was originally developed for an innovative information exploration tool, called Ariadne, which operates on bibliographic databases with tens of millions of records. In this paper, the method behind Ariadne is further developed and applied to the research question of the special issue "Same data, different results" - the better understanding of topic (re-)construction by different bibliometric approaches. For the case of the Astro dataset of 111,616 articles in astronomy and astrophysics, a new instantiation of the interactive exploring tool, LittleAriadne, has been created. This paper contributes to the overall challenge to delineate and define topics in two different ways. First, we produce two clustering solutions based on vector representations of articles in a lexical space. These vectors are built on semantic indexing of entities associated with those articles. Second, we discuss how LittleAriadne can be used to browse through the network of topical terms, authors, journals, citations and various cluster solutions of the Astro dataset. More specifically, we treat the assignment of an article to the different clustering solutions as an additional element of its bibliographic record. Keeping the principle of semantic indexing on the level of such an extended list of entities of the bibliographic record, LittleAriadne in turn provides a visualization of the context of a specific clustering solution. It also conveys the similarity of article clusters produced by different algorithms, hence representing a complementary approach to other possible means of comparison.

cs.DL↗

Bibliometrics and Information Retrieval: Creating Knowledge through Research Synergies

This panel brings together experts in bibliometrics and information retrieval to discuss how each of these two important areas of information science can help to inform the research of the other. There is a growing body of literature that capitalizes on the synergies created by combining methodological approaches of each to solve research problems and practical issues related to how information is created, stored, organized, retrieved and used. The session will begin with an overview of the common threads that exist between IR and metrics, followed by a summary of findings from the BIR workshops and examples of research projects that combine aspects of each area to benefit IR or metrics research areas, including search results ranking, semantic indexing and visualization. The panel will conclude with an engaging discussion with the audience to identify future areas of research and collaboration.

cs.IR↗