SearcharxivSearch

arXiv subjects

Lars Vogt

Publications and source records attributed to Lars Vogt.

12 recordsLinked to original sources

The TripleA Principle: Making Knowledge Actionable, Applicable, and Auditable in Post-FAIR Infrastructures

Ecological restoration, species distribution modelling, and invasive species management share a difficulty: knowledge that is findable and reusable carries no explicit account of the conditions under which it can be validly applied, or of the evidence grounding them. Applying it correctly is therefore demanding and expert-dependent, and misapplication usually goes unrecorded. The FAIR and CLEAR principles improved the findability, accessibility, interoperability, reusability, and human-interpretability of knowledge, but these address properties of representation, and reliable action requires more. Bridging the knowledge-action gap requires characterizing knowledge in terms of the operations it supports. Analysing what an operation needs, we derive three capabilities a knowledge representation must support. Actionability is the capacity to supply the knowledge and objective an operation executes. Applicability is the capacity to assess whether it can be reliably performed, through explicit conditions evaluated against context. Auditability is the capacity to assess the empirical grounding for that reliability, through documented success and failure. These form the three criteria of the TripleA Principle, an implementation-indipendent guide for next-generation knowledge infrastructures, jointly sufficient for the representational preconditions of reliably grounded action though not for its justification. Building on the Semantic Units Framework, we realize the principle as action units, typed components in which the knowledge an operation executes, the conditions under which it may validly be applied, and its documented successes and failures are addressable and evaluable. Action units form a nested hierarchy in which documented failure refines the conditions of valid use, letting knowledge graphs act as context-sensitive, evidentially accountable decision-support systems.

cs.DB

The Semantic Ladder: A Framework for Progressive Formalization of Natural Language Content for Knowledge Graphs and AI Systems

Semantic data and knowledge infrastructures must reconcile two fundamentally different forms of representation: natural language, in which most knowledge is created and communicated, and formal semantic models, which enable machine-actionable integration, interoperability, and reasoning. Bridging this gap remains a central challenge, particularly when full semantic formalization is required at the point of data entry. Here, we introduce the Semantic Ladder, an architectural framework that enables the progressive formalization of data and knowledge. Building on the concept of modular semantic units as identifiable carriers of meaning, the framework organizes representations across levels of increasing semantic explicitness, ranging from natural language text snippets to ontology-based and higher-order logical models. Transformations between levels support semantic enrichment, statement structuring, and logical modelling while preserving semantic continuity and traceability. This approach enables the incremental construction of semantic knowledge spaces, reduces the semantic parsing burden, and supports the integration of heterogeneous representations, including natural language, structured semantic models, and vector-based embeddings. The Semantic Ladder thereby provides a foundation for scalable, interoperable, and AI-ready data and knowledge infrastructures.

cs.CL

The Grammar of FAIR: A Granular Architecture of Semantic Units for FAIR Semantics, Inspired by Biology and Linguistics

The FAIR Principles aim to make data and knowledge Findable, Accessible, Interoperable, and Reusable, yet current digital infrastructures often lack a unifying semantic framework that bridges human cognition and machine-actionability. In this paper, we introduce the Grammar of FAIR: a granular and modular architecture for FAIR semantics built on the concept of semantic units. Semantic units, comprising atomic statement units and composite compound units, implement the principle of semantic modularisation, decomposing data and knowledge into independently identifiable, semantically meaningful, and machine-actionable units. A central metaphor guiding our approach is the analogy between the hierarchy of level of organisation in biological systems and the hierarchy of levels of organisation in information systems: both are structured by granular building blocks that mediate across multiple perspectives while preserving functional unity. Drawing further inspiration from concept formation and natural language grammar, we show how these building blocks map to FAIR Digitial Objects (FDOs), enabling format-agnostic semantic transitivity from natural language token models to schema-based representations. This dual biological-linguistic analogy provides a semantics-first foundation for evolving cross-ecosystem infrastructures, paving the way for the Internet of FAIR Data and Services (IFDS) and a future of modular, AI-ready, and citation-granular scholarly communication.

cs.DB

A Framework for FAIR and CLEAR Ecological Data and Knowledge: Semantic Units for Synthesis and Causal Modelling

Ecological research increasingly relies on integrating heterogeneous datasets and knowledge to explain and predict complex phenomena. Yet, differences in data types, terminology, and documentation often hinder interoperability, reuse, and causal understanding. We present the Semantic Units Framework, a novel, domain-agnostic semantic modelling approach applied here to ecological data and knowledge in compliance with the FAIR (Findable, Accessible, Interoperable, Reusable) and CLEAR (Cognitively interoperable, semantically Linked, contextually Explorable, easily Accessible, human-Readable and -interpretable) Principles. The framework models data and knowledge as modular, logic-aware semantic units: single propositions (statement units) or coherent groups of propositions (compound units). Statement units can model measurements, observations, or universal relationships, including causal ones, and link to methods and evidence. Compound units group related statement units into reusable, semantically coherent knowledge objects. Implemented using RDF, OWL, and knowledge graphs, semantic units can be serialized as FAIR Digital Objects with persistent identifiers, provenance, and semantic interoperability. We show how universal statement units build ecological causal networks, which can be composed into causal maps and perspective-specific subnetworks. These support causal reasoning, confounder detection (back-door), effect identification with unobserved confounders (front-door), application of do-calculus, and alignment with Bayesian networks, structural equation models, and structural causal models. By linking fine-grained empirical data to high-level causal reasoning, the Semantic Units Framework provides a foundation for ecological knowledge synthesis, evidence annotation, cross-domain integration, reproducible workflows, and AI-ready ecological research.

cs.DB

Rosetta Statements: Simplifying FAIR Knowledge Graph Construction with a User-Centered Approach

Machines need data and metadata to be machine-actionable and FAIR (findable, accessible, interoperable, reusable) to manage increasing data volumes. Knowledge graphs and ontologies are key to this, but their use is hampered by high access barriers due to required prior knowledge in semantics and data modelling. The Rosetta Statement approach proposes modeling English natural language statements instead of a mind-independent reality. We propose a metamodel for creating semantic schema patterns for simple statement types. The approach supports versioning of statements and provides a detailed editing history. Each Rosetta Statement pattern has a dynamic label for displaying statements as natural language sentences. Implemented in the Open Research Knowledge Graph (ORKG) as a use case, this approach allows domain experts to define data schema patterns without needing semantic knowledge. Future plans include combining Rosetta Statements with semantic units to organize ORKG into meaningful subgraphs, improving usability. A search interface for querying statements without needing SPARQL or Cypher knowledge is also planned, along with tools for data entry and display using Large Language Models. The Rosetta Statement metamodel supports a three-step knowledge graph construction procedure. Domain experts can model semantic content without support from ontology engineers by using Wikidata, lowering entry barriers and increasing cognitive interoperability. The second level involves mapping Wikidata terms to established ontologies, and the third step developing semantic graph patterns for reasoning, requiring collaboration with ontology engineers.

cs.DB

Rethinking OWL Expressivity: Semantic Units for FAIR and Cognitively Interoperable Knowledge Graphs Why OWLs don't have to understand everything they say

Semantic knowledge graphs are foundational to implementing the FAIR Principles, yet RDF/OWL representations often lack the semantic flexibility and cognitive interoperability required in scientific domains. We present a novel framework for semantic modularization based on semantic units (i.e., modular, semantically coherent subgraphs enhancing expressivity, reusability, and interpretability), combined with four new representational resource types (some-instance, most-instances, every-instance, all-instances) for modelling assertional, contingent, prototypical, and universal statements. The framework enables the integration of knowledge modelled using different logical frameworks (e.g., OWL, First-Order Logic, or none), provided each semantic unit is internally consistent and annotated with its logic base. This allows, for example, querying all OWL 2.0-compliant units for reasoning purposes while preserving the full graph for broader knowledge discovery. Our framework addresses twelve core limitations of OWL/RDF modeling, including negation, cardinality, complex class axioms, conditional and directive statements, and logical arguments, while improving cognitive accessibility for domain experts. We provide schemata and translation patterns to demonstrate semantic interoperability and reasoning potential, establishing a scalable foundation for constructing FAIR-aligned, semantically rich knowledge graphs.

cs.DB

FAIR 2.0: Extending the FAIR Guiding Principles to Address Semantic Interoperability

FAIR data presupposes their successful communication between machines and humans while preserving their meaning and reference, requiring all parties involved to share the same background knowledge. Inspired by English as a natural language, we investigate the linguistic structure that ensures reliable communication of information and draw parallels with data structures, understanding both as models of systems of interest. We conceptualize semantic interoperability as comprising terminological and propositional interoperability. The former includes ontological (i.e., same meaning) and referential (i.e., same referent/extension) interoperability and the latter schema (i.e., same data schema) and logical (i.e., same logical framework) interoperability. Since no best ontology and no best data schema exists, establishing semantic interoperability and FAIRness of data and metadata requires the provision of a comprehensive set of relevant ontological and referential entity mappings and schema crosswalks. We therefore propose appropriate additions to the FAIR Guiding Principles, leading to FAIR 2.0. Furthermore, achieving FAIRness of data requires the provision of FAIR services in addition to organizing data into FAIR Digital Objects. FAIR services include a terminology, a schema, and an operations service.

cs.DB

FAIR Knowledge Graphs with Semantic Units: a Prototype

Knowledge graphs and ontologies are becoming increasingly important in the context of making data and metadata findable, accessible, interoperable, and reusable (FAIR). We introduce the concept of Semantic Units for organizing Knowledge Graphs into identifiable and semantically meaningful subgraphs. Each Semantic Unit is represented in the graph by its own resource that instantiates a Semantic Unit class. Different types of Semantic Units are distinguished, and together they can organize a Knowledge Graph into different levels of representational granularity with partially overlapping, partially enclosed subgraphs that users of Knowledge Graphs can refer to for making statements about statements. The use of Semantic Units in Knowledge Graphs supports making them FAIR and increases the human-reader-actionability of their data and metadata by increasing the graph's cognitive interoperability by increasing its explorability for a human reader. We introduce a minimal prototype web application for a user-driven FAIR Knowledge Graph that is based on Semantic Units.

cs.DB

Towards a Rosetta Stone for (meta)data: Learning from natural language to improve semantic and cognitive interoperability

In order to effectively manage the overwhelming influx of data, it is crucial to ensure that data is findable, accessible, interoperable, and reusable (FAIR). While ontologies and knowledge graphs have been employed to enhance FAIRness, challenges remain regarding semantic and cognitive interoperability. We explore how English facilitates reliable communication of terms and statements, and transfer our findings to a framework of ontologies and knowledge graphs, while treating terms and statements as minimal information units. We categorize statement types based on their predicates, recognizing the limitations of modeling non-binary predicates with multiple triples, which negatively impacts interoperability. Terms are associated with different frames of reference, and different operations require different schemata. Term mappings and schema crosswalks are therefore vital for semantic interoperability. We propose a machine-actionable Rosetta Stone Framework for (meta)data, which uses reference terms and schemata as an interlingua to minimize mappings and crosswalks. Modeling statements rather than a human-independent reality ensures cognitive familiarity and thus better interoperability of data structures. We extend this Rosetta modeling paradigm to reference schemata, resulting in simple schemata with a consistent structure across statement types, empowering domain experts to create their own schemata using the Rosetta Editor, without requiring knowledge of semantics. The Editor also allows specifying textual and graphical display templates for each schema, delivering human-readable data representations alongside machine-actionable data structures. The Rosetta Query Builder derives queries based on completed input forms and the information from corresponding reference schemata. This work sets the conceptual ground for the Rosetta Stone Framework that we plan to develop in the future.

cs.DB

Knowledge Graph Building Blocks: An easy-to-use Framework for developing FAIREr Knowledge Graphs

Knowledge graphs and ontologies provide promising technical solutions for implementing the FAIR Principles for Findable, Accessible, Interoperable, and Reusable data and metadata. However, they also come with their own challenges. Nine such challenges are discussed and associated with the criterion of cognitive interoperability and specific FAIREr principles (FAIR + Explorability raised) that they fail to meet. We introduce an easy-to-use, open source knowledge graph framework that is based on knowledge graph building blocks (KGBBs). KGBBs are small information modules for knowledge-processing, each based on a specific type of semantic unit. By interrelating several KGBBs, one can specify a KGBB-driven FAIREr knowledge graph. Besides implementing semantic units, the KGBB Framework clearly distinguishes and decouples an internal in-memory data model from data storage, data display, and data access/export models. We argue that this decoupling is essential for solving many problems of knowledge management systems. We discuss the architecture of the KGBB Framework as we envision it, comprising (i) an openly accessible KGBB-Repository for different types of KGBBs, (ii) a KGBB-Engine for managing and operating FAIREr knowledge graphs (including automatic provenance tracking, editing changelog, and versioning of semantic units); (iii) a repository for KGBB-Functions; (iv) a low-code KGBB-Editor with which domain experts can create new KGBBs and specify their own FAIREr knowledge graph without having to think about semantic modelling. We conclude with discussing the nine challenges and how the KGBB Framework provides solutions for the issues they raise. While most of what we discuss here is entirely conceptual, we can point to two prototypes that demonstrate the principle feasibility of using semantic units and KGBBs to manage and structure knowledge graphs.

cs.DB

The FAIREr Guiding Principles: Organizing data and metadata into semantically meaningful types of FAIR Digital Objects to increase their human explorability and cognitive interoperability

Ensuring the FAIRness (Findable, Accessible, Interoperable, Reusable) of data and metadata is an important goal in both research and industry. Knowledge graphs and ontologies have been central in achieving this goal, with interoperability receiving much attention. The European Open Science Cloud initiative distinguishes in its Interoperability Framework (EOSC IF) four layers of interoperability: technical, semantic, organizational, and legal. This paper argues that the emphasis on machine-actionability has overshadowed the essential need for human-actionability of data and metadata, and provides three examples that describe the lack of human-actionability within knowledge graphs. In response, the paper propagates the incorporation of cognitive interoperability as another vital layer within the EOSC IF and discusses the relation between human explorability of data and metadata and their cognitive interoperability. It suggests adding the Principle of human Explorability, extending FAIR to the FAIREr Guiding Principles. The subsequent sections present the concept of semantic units, elucidating their important role in enhancing the human explorability and cognitive interoperability of knowledge graphs. Each semantic unit can be displayed in a user interface either as a mind-map-like graph or as natural language text. Semantic units organize knowledge graphs into levels of representational granularity, distinct granularity trees, and diverse frames of reference. Semantic units can be represented as FAIR Digital Objects (FDOs) that instantiate corresponding FDO classes. Various categories of FDOs are distinguished. The development of innovative user interfaces enabled by FDOs that are based on semantic units would empower users to access, navigate, and explore information in knowledge graphs with enhanced efficiency.

cs.DB

Semantic Units: Organizing knowledge graphs into semantically meaningful units of representation

Knowledge graphs and ontologies are becoming increasingly important as technical solutions for Findable, Accessible, Interoperable, and Reusable data and metadata (FAIR Guiding Principles). We discuss four challenges that impede the use of FAIR knowledge graphs and propose semantic units as their potential solution. Semantic units structure a knowledge graph into identifiable and semantically meaningful subgraphs. Each unit is represented by its own resource, instantiates a corresponding semantic unit class, and can be implemented as a FAIR Digital Object and a nanopublication in RDF/OWL and property graphs. We distinguish statement and compound units as basic categories of semantic units. Statement units represent smallest, independent propositions that are semantically meaningful for a human reader. They consist of one or more triples and mathematically partition a knowledge graph. We distinguish assertional, contingent (prototypical), and universal statement units as basic types of statement units and propose representational schemes and formal semantics for them (including for absence statements, negations, and cardinality restrictions) that do not involve blank nodes and that translate back to OWL. Compound units, on the other hand, represent semantically meaningful collections of semantic units and we distinguish various types of compound units, representing different levels of representational granularity, different types of granularity trees, and different frames of reference. Semantic units support making statements about statements, can be used for graph-alignment, subgraph-matching, knowledge graph profiling, and for managing access restrictions to sensitive data. Organizing the graph into semantic units supports the separation of ontological, diagnostic (i.e., referential), and discursive information, and it also supports the differentiation of multiple frames of reference.

cs.DB