Searcharxiv⌕ Search

arXiv subjects

Markus Schröder

Publications and source records attributed to Markus Schröder.

At least 19 recordsLinked to original sources

How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance

When using Large Language Models (LLMs) to support Knowledge Graph Engineering (KGE), one of the first indications when searching for an appropriate model is its size. According to the scaling laws, larger models typically show higher capabilities. However, in practice, resource costs are also an important factor and thus it makes sense to consider the ratio between model performance and costs. The LLM-KG-Bench framework enables the comparison of LLMs in the context of KGE tasks and assesses their capabilities of understanding and producing KGs and KG queries. Based on a dataset created in an LLM-KG-Bench run covering 26 open state-of-the-art LLMs, we explore the model size scaling laws specific to KGE tasks. In our analyses, we assess how benchmark scores evolve between different model size categories. Additionally, we inspect how the general score development of single models and families of models correlates to their size. Our analyses revealed that, with a few exceptions, the model size scaling laws generally also apply to the selected KGE tasks. However, in some cases, plateau or ceiling effects occurred, i.e., the task performance did not change much between a model and the next larger model. In these cases, smaller models could be considered to achieve high cost-effectiveness. Regarding models of the same family, sometimes larger models performed worse than smaller models of the same family. These effects occurred only locally. Hence it is advisable to additionally test the next smallest and largest model of the same family.

cs.AI↗

Full Quantum dynamics study for H atom scattering from graphene

This study deals with the understanding of hydrogen atom scattering from graphene, a process critical for exploring C-H bond formation and energy transfer during the atom surface collision. In our previous work (J.Chem.Phys \textbf{159}, 194102, (2023)), starting from a cell with 24 carbon atoms treated periodically, we have achieved quantum dynamics (QD) simulations with a reduced-dimensional model (15D) and a simulation in full dimensionality (75D). In the former work, the H atom attacked the top of a single C atom, enabling a comparison of QD simulation results with classical molecular dynamics (cMD). Our approach required the use of sophisticated techniques such as Monte Carlo Canonical Polyadic Decomposition (MCCPD) and Multilayer Multi-Configuration Time-Dependent Hartree (ML-MCTDH), as well as a further development of quantum flux calculations. We could benchmark our calculations by comparison with cMD calculations. We have now refined our method to better mimic experimental conditions. Specifically, rather than sending the H atom to a specific position on the surface, we have employed a plane wave for the H atom in directions parallel to the surface. Key findings for these new simulations include the identification of discrepancies between classical molecular dynamics (cMD) simulations and experiments, which are attributed to both the potential energy surface (PES) and quantum effects. Additionally, the study sheds light on the role of classical collective normal modes during collisions, providing insights into energy transfer processes. The results validate the robustness of our simulation methodologies and highlight the importance of considering quantum mechanical effects in the study of hydrogen-graphene interactions.

physics.chem-ph↗

Data Collection of Real-Life Knowledge Work in Context: The RLKWiC Dataset

Over the years, various approaches have been employed to enhance the productivity of knowledge workers, from addressing psychological well-being to the development of personal knowledge assistants. A significant challenge in this research area has been the absence of a comprehensive, publicly accessible dataset that mirrors real-world knowledge work. Although a handful of datasets exist, many are restricted in access or lack vital information dimensions, complicating meaningful comparison and benchmarking in the domain. This paper presents RLKWiC, a novel dataset of Real-Life Knowledge Work in Context, derived from monitoring the computer interactions of eight participants over a span of two months. As the first publicly available dataset offering a wealth of essential information dimensions (such as explicated contexts, textual contents, and semantics), RLKWiC seeks to address the research gap in the personal information management domain, providing valuable insights for modeling user behavior.

cs.AI↗

CO-Fun: A German Dataset on Company Outsourcing in Fund Prospectuses for Named Entity Recognition and Relation Extraction

The process of cyber mapping gives insights in relationships among financial entities and service providers. Centered around the outsourcing practices of companies within fund prospectuses in Germany, we introduce a dataset specifically designed for named entity recognition and relation extraction tasks. The labeling process on 948 sentences was carried out by three experts which yields to 5,969 annotations for four entity types (Outsourcing, Company, Location and Software) and 4,102 relation annotations (Outsourcing-Company, Company-Location). State-of-the-art deep learning models were trained to recognize entities and extract relations showing first promising results. An anonymized version of the dataset, along with guidelines and the code used for model training, are publicly available at https://www.dfki.uni-kl.de/cybermapping/data/CO-Fun-1.0-anonymized.zip.

cs.CL↗

Compact sum-of-products form of the molecular electronic Hamiltonian based on canonical polyadic decomposition

We propose an approach to represent the second-quantized electronic Hamiltonian in a compact sum-of-products (SOP) form. The approach is based on the canonical polyadic decomposition (CPD) of the original Hamiltonian projected onto the sub-Fock spaces formed by groups of spin orbitals. The algorithm for obtaining the canonical polyadic form starts from an exact sum-of-products, which is then optimally compactified using an alternating least-squares procedure. We discuss the relation of this specific SOP with related forms, namely the Tucker format and the matrix product operator often used in conjunction with matrix product states. We benchmark the method on the electronic dynamics of an excited water molecule, trans-polyenes, and the charge migration in glycine upon inner-valence ionization. The quantum dynamics are performed with the multilayer multi-configuration time-dependent Hartree method in second quantization representation (MCTDH-SQR). Other methods based on tree-tensor Ansätze may profit from this general approach.

physics.chem-ph↗

Towards Self-organizing Personal Knowledge Assistants in Evolving Corporate Memories

This paper presents a retrospective overview of a decade of research in our department towards self-organizing personal knowledge assistants in evolving corporate memories. Our research is typically inspired by real-world problems and often conducted in interdisciplinary collaborations with research and industry partners. We summarize past experiments and results comprising topics like various ways of knowledge graph construction in corporate and personal settings, Managed Forgetting and (Self-organizing) Context Spaces as a novel approach to Personal Information Management (PIM) and knowledge work support. Past results are complemented by an overview of related work and some of our latest findings not published so far. Last, we give an overview of our related industry use cases including a detailed look into CoMem, a Corporate Memory based on our presented research already in productive use and providing challenges for further research. Many contributions are only first steps in new directions with still a lot of untapped potential, especially with regard to further increasing the automation in PIM and knowledge work support.

cs.AI↗

State-resolved infrared spectrum of the protonated water dimer: Revisiting the characteristic proton transfer doublet peak

The infrared (IR) spectra of protonated water clusters encode precise information on the dynamics and structure of the hydrated proton. However, the strong anharmonic coupling and quantum effects of these elusive species remain puzzling up to the present day. Here, we report unequivocal evidence that the interplay between the proton transfer and the water wagging motions in the protonated water dimer (Zundel ion) giving rise to the characteristic doublet peak is both more complex and more sensitive to subtle energetic changes than previously thought. In particular, hitherto overlooked low-intensity satellite peaks in the experimental spectrum are now unveiled and mechanistically assigned. Our findings rely on the comparison of IR spectra obtained using two highly accurate potential energy surfaces in conjunction with highly accurate state-resolved quantum simulations. We demonstrate that these high-accuracy simulations are important for providing definite assignments of the complex IR signals of fluxional molecules.

physics.chem-ph↗

Functional Component Descriptions for Electrical Circuits based on Semantic Technology Reasoning

Circuit diagrams have been used in electrical engineering for decades to describe the wiring of devices and facilities. They depict electrical components in a symbolic and graph-based manner. While the circuit design is usually performed electronically, there are still legacy paper-based diagrams that require digitization in order to be used in CAE systems. Generally, knowledge on specific circuits may be lost between engineering projects, making it hard for domain novices to understand a given circuit design. The graph-based nature of these documents can be exploited by semantic technology-based reasoning in order to generate human-understandable descriptions of their functional principles. More precisely, each electrical component (e.g. a diode) of a circuit may be assigned a high-level function label which describes its purpose within the device (e.g. flyback diode for reverse voltage protection). In this paper, forward chaining rules are used for such a generation. The described approach is applicable for both CAE-based circuits as well as raw circuits yielded by an image understanding pipeline. The viability of the approach is demonstrated by application to an existing set of circuits.

cs.OH↗

The coupling of the hydrated proton to its first solvation shell

The transfer of a hydrated proton between water molecules in aqueous solution is accompanied by the large-scale structural reorganization of the environment as the proton relocates, giving rise to the Grotthus mechanism. The Zundel (H5O2+) and Eigen (H9O4+) cations are the main intermediate structures in this process. They exhibit radically different gas-phase infrared (IR) spectra, indicating fundamentally different environments of the solvated proton in its first solvation shell. The question arises: is there a least common denominator structure that explains the IR spectra of the Zundel and Eigen cations, and hence of the solvated proton? Full dimensional quantum simulations of these protonated cations demonstrate that two dynamical water molecules embedded in the static environment of the parent Eigen cation constitute this fundamental subunit. It is sufficient to explain the spectral signatures and anharmonic couplings of the solvated proton in its first solvation shell. In particular, we identify the anharmonic vibrational modes that explain the large broadening of the proton transfer peak in the experimental IR spectrum of the Eigen cation, of which the origin remained so far unclear. Our findings about the quantum mechanical structure of the first solvation shell provide a starting point for further investigations of the larger protonated water clusters with second and additional solvation shells.

physics.chem-ph↗

Spread2RML: Constructing Knowledge Graphs by Predicting RML Mappings on Messy Spreadsheets

The RDF Mapping Language (RML) allows to map semi-structured data to RDF knowledge graphs. Besides CSV, JSON and XML, this also includes the mapping of spreadsheet tables. Since spreadsheets have a complex data model and can become rather messy, their mapping creation tends to be very time consuming. In order to reduce such efforts, this paper presents Spread2RML which predicts RML mappings on messy spreadsheets. This is done with an extensible set of RML object map templates which are applied for each column based on heuristics. In our evaluation, three datasets are used ranging from very messy synthetic data to spreadsheets from data.gov which are less messy. We obtained first promising results especially with regard to our approach being fully automatic and dealing with rather messy data.

cs.DB↗

Mapping Spreadsheets to RDF: Supporting Excel in RML

The RDF Mapping Language (RML) enables, among other formats, the mapping of tabular data as Comma-Separated Values (CSV) files to RDF graphs. Unfortunately, the widely used spreadsheet format is currently neglected by its specification and well-known implementations. Therefore, we extended one of the tools which is RML Mapper to support Microsoft Excel spreadsheet files and demonstrate its capabilities in an interactive online demo. Our approach allows to access various meta data of spreadsheet cells in typical RML maps. Some experimental features for more specific use cases are also provided. The implementation code is publicly available in a GitHub fork.

cs.DB↗

A Linked Data Application Framework to Enable Rapid Prototyping

Application developers, in our experience, tend to hesitate when dealing with linked data technologies. To reduce their initial hurdle and enable rapid prototyping, we propose in this paper a framework for building linked data applications. Our approach especially considers the participation of web developers and non-technical users without much prior knowledge about linked data concepts. Web developers are supported with bidirectional RDF to JSON conversions and suitable CRUD endpoints. Non-technical users can browse websites generated from JSON data by means of a template language. A prototypical open source implementation demonstrates its capabilities.

cs.DB↗

Dataset Generation Patterns for Evaluating Knowledge Graph Construction

Confidentiality hinders the publication of authentic, labeled datasets of personal and enterprise data, although they could be useful for evaluating knowledge graph construction approaches in industrial scenarios. Therefore, our plan is to synthetically generate such data in a way that it appears as authentic as possible. Based on our assumption that knowledge workers have certain habits when they produce or manage data, generation patterns could be discovered which can be utilized by data generators to imitate real datasets. In this paper, we initially derived 11 distinct patterns found in real spreadsheets from industry and demonstrate a suitable generator called Data Sprout that is able to reproduce them. We describe how the generator produces spreadsheets in general and what altering effects the implemented patterns have.

cs.DB↗

Interactively Constructing Knowledge Graphs from Messy User-Generated Spreadsheets

When spreadsheets are filled freely by knowledge workers, they can contain rather unstructured content. For humans and especially machines it becomes difficult to interpret such data properly. Therefore, spreadsheets are often converted to a more explicit, formal and structured form, for example, to a knowledge graph. However, if a data maintenance strategy has been missing and user-generated data becomes "messy", the construction of knowledge graphs will be a challenging task. In this paper, we catalog several of those challenges and propose an interactive approach to solve them. Our approach includes a graphical user interface which enables knowledge engineers to bulk-annotate spreadsheet cells with extracted information. Based on the cells' annotations a knowledge graph is ultimately formed. Using five spreadsheets from an industrial scenario, we built a 25k-triple graph during our evaluation. We compared our method with the state-of-the-art RDF Mapping Language (RML) attempt. The comparison highlights contributions of our approach.

cs.DB↗

Bridging the Technology Gap Between Industry and Semantic Web: Generating Databases and Server Code From RDF

Despite great advances in the area of Semantic Web, industry rather seldom adopts Semantic Web technologies and their storage and query concepts. Instead, relational databases (RDB) are often deployed to store business-critical data, which are accessed via REST interfaces. Yet, some enterprises would greatly benefit from Semantic Web related datasets which are usually represented with the Resource Description Framework (RDF). To bridge this technology gap, we propose a fully automatic approach that generates suitable RDB models with REST APIs to access them. In our evaluation, generated databases from different RDF datasets are examined and compared. Our findings show that the databases sufficiently reflect their counterparts while the API is able to reproduce rather simple SPARQL queries. Potentials for improvements are identified, for example, the reduction of data redundancies in generated databases.

cs.DB↗

The Person Index Challenge: Extraction of Persons from Messy, Short Texts

When persons are mentioned in texts with their first name, last name and/or middle names, there can be a high variation which of their names are used, how their names are ordered and if their names are abbreviated. If multiple persons are mentioned consecutively in very different ways, especially short texts can be perceived as "messy". Once ambiguous names occur, associations to persons may not be inferred correctly. Despite these eventualities, in this paper we ask how well an unsupervised algorithm can build a person index from short texts. We define a person index as a structured table that distinctly catalogs individuals by their names. First, we give a formal definition of the problem and describe a procedure to generate ground truth data for future evaluations. To give a first solution to this challenge, a baseline approach is implemented. By using our proposed evaluation strategy, we test the performance of the baseline and suggest further improvements. For future research the source code is publicly available.

cs.CL↗

Data Preparation in Agriculture Through Automated Semantic Annotation -- Basis for a Wide Range of Smart Services

Modern agricultural technology and the increasing digitalisation of such processes provide a wide range of data. However, their efficient and beneficial use suffers from legitimate concerns about data sovereignty and control, format inconsistencies and different interpretations. As a proposed solution, we present Wikinormia, a collaborative platform in which interested participants can describe and discuss their own new data formats. Once a finalized vocabulary has been created, specific parsers can semantically process the raw data into three basic representations: spatial information, time series and semantic facts (agricultural knowledge graph). Thanks to publicly accessible definitions and descriptions, developers can easily gain an overview of the concepts that are relevant to them. A variety of services will then (subject to individual access rights) be able to query their data simply via a query interface and retrieve results. We have already implemented this proposed solution in a prototype in the SDSD (Smart Data - Smart Services) project and demonstrate the benefits with a range of representative services. This provides an efficient system for the cooperative, flexible digitalisation of agricultural workflows.

cs.CY↗

Interactive Concept Mining on Personal Data -- Bootstrapping Semantic Services

Semantic services (e.g. Semantic Desktops) are still afflicted by a cold start problem: in the beginning, the user's personal information sphere, i.e. files, mails, bookmarks, etc., is not represented by the system. Information extraction tools used to kick-start the system typically create 1:1 representations of the different information items. Higher level concepts, for example found in file names, mail subjects or in the content body of these items, are not extracted. Leaving these concepts out may lead to underperformance, having to many of them (e.g. by making every found term a concept) will clutter the arising knowledge graph with non-helpful relations. In this paper, we present an interactive concept mining approach proposing concept candidates gathered by exploiting given schemata of usual personal information management applications and analysing the personal information sphere using various metrics. To heed the subjective view of the user, a graphical user interface allows to easily rank and give feedback on proposed concept candidates, thus keeping only those actually considered relevant. A prototypical implementation demonstrates major steps of our approach.

cs.CL↗