SearcharxivSearch

arXiv subjects

Andrew Mitchell

Publications and source records attributed to Andrew Mitchell.

At least 19 recordsLinked to original sources

The AnIML Ontology: Enabling Semantic Interoperability for Large-Scale Experimental Data in Interconnected Scientific Labs

Achieving semantic interoperability across heterogeneous experimental data systems remains a major barrier to data-driven scientific discovery. The Analytical Information Markup Language (AnIML), a flexible XML-based standard for analytical chemistry and biology, is increasingly used in industrial R&D labs for managing and exchanging experimental data. However, the expressivity of the XML schema permits divergent interpretations across stakeholders, introducing inconsistencies that undermine the interoperability the AnIML schema was designed to support. In this paper, we present the AnIML Ontology, an OWL 2 ontology that formalises the semantics of AnIML and aligns it with the Allotrope Data Format to support future cross-system and cross-lab interoperability. The ontology was developed using an expert-in-the-loop approach combining LLM-assisted requirement elicitation with collaborative ontology engineering. We validate the ontology through a multi-layered approach: data-driven transformation of real-world AnIML files into knowledge graphs, competency question verification via SPARQL, and a novel validation protocol based on adversarial negative competency questions mapped to established ontological anti-patterns and enforced via SHACL constraints.

cs.AI

IDEA2: Expert-in-the-loop competency question elicitation for collaborative ontology engineering

Competency question (CQ) elicitation represents a critical but resource-intensive bottleneck in ontology engineering. This foundational phase is often hampered by the communication gap between domain experts, who possess the necessary knowledge, and ontology engineers, who formalise it. This paper introduces IDEA2, a novel, semi-automated workflow that integrates Large Language Models (LLMs) within a collaborative, expert-in-the-loop process to address this challenge. The methodology is characterised by a core iterative loop: an initial LLM-based extraction of CQs from requirement documents, a co-creational review and feedback phase by domain experts on an accessible collaborative platform, and an iterative, feedback-driven reformulation of rejected CQs by an LLM until consensus is achieved. To ensure transparency and reproducibility, the entire lifecycle of each CQ is tracked using a provenance model that captures the full lineage of edits, anonymised feedback, and generation parameters. The workflow was validated in 2 real-world scenarios (scientific data, cultural heritage), demonstrating that IDEA2 can accelerate the requirements engineering process, improve the acceptance and relevance of the resulting CQs, and exhibit high usability and effectiveness among domain experts. We release all code and experiments at https://github.com/KE-UniLiv/IDEA2

cs.AI

Disentangling the schema turn: Restoring the information base to conceptual modelling

If one looks at contemporary mainstream development practices for conceptual modelling in computer science, these so clearly focus on a conceptual schema completely separated from its information base that the conceptual schema is often just called the conceptual model. These schema-centric practices are crystallized in almost every database textbook. We call this strong, almost universal, bias towards conceptual schemas the schema turn. The focus of this paper is on disentangling this turn within (computer science) conceptual modeling. It aims to shed some light on how it emerged and so show that it is not fundamental. To show that modern technology enables the adoption of an inclusive schema-and-base conceptual modelling approach, which in turn enables more automated, and empirically motivated practices. And to show, more generally, the space of possible conceptual modelling practices is wider than currently assumed. It also uses the example of bCLEARer to show that the implementations in this wider space will probably need to rely on new pipeline-based conceptual modelling techniques. So, it is possible that the schema turn's complete exclusion of the information base could be merely a temporary evolutionary detour.

cs.DB

The matrix potential game and structures of self-affine sets

We present a new variant of the potential game and show that certain compact subsets of $\mathbb{R}^n$, including a large class of self-affine sets, are winning in our game. We prove that sets with sufficiently strong winning conditions are non-empty, provide a lower bound for their Hausdorff dimension, show that they have good intersection properties, and provide conditions under which, given $M \in \mathbb{N}$, they contain a homothetic copy of every set with at most $M$ elements. The applications of our game to self-affine sets are new and complement the recent work of Yavicoli et al (Math. Z. 2022 and Int. Math. Res. Not. IMRN 2023) for self-similar sets.

math.DS

Digitalizing Uncertain Information

The paper sketches some initial results from an ongoing project to develop an ontology-based digital form for representing uncertain information. We frame this work as a journey from lower to higher levels of digital maturity across a technology divide. The paper first sets a baseline by describing the basic challenges any project dealing with digital uncertainty faces. It then describes how the project is facing them. It shows firstly how an extensional ontology (such as the BORO Foundational Ontology or the Information Exchange Standard) can be extended with a Lewisian counterpart approach to formalizing uncertainty that is adapted to computing. And then it shows how this is expressive enough to handle the challenges. Keywords: actuality, BORO Foundational Ontology, counterpart, Information Exchange Standard, informational uncertainty, my doxastic actualities, two-dimensional semantics.

cs.DB

RELRaE: LLM-Based Relationship Extraction, Labelling, Refinement, and Evaluation

A large volume of XML data is produced in experiments carried out by robots in laboratories. In order to support the interoperability of data between labs, there is a motivation to translate the XML data into a knowledge graph. A key stage of this process is the enrichment of the XML schema to lay the foundation of an ontology schema. To achieve this, we present the RELRaE framework, a framework that employs large language models in different stages to extract and accurately label the relationships implicitly present in the XML schema. We investigate the capability of LLMs to accurately generate these labels and then evaluate them. Our work demonstrates that LLMs can be effectively used to support the generation of relationship labels in the context of lab automation, and that they can play a valuable role within semi-automatic ontology generation frameworks more generally.

cs.AI

Diffraction of the Hat and Spectre tilings and some of their relatives

The diffraction spectra of the Hat and Spectre monotile tilings, which are known to be pure point, are derived and computed explicitly. This is done via model set representatives of self-similar members in the topological conjugacy classes of the Hat and the Spectre tiling, which are the CAP and the CASPr tiling, respectively. This is followed by suitable reprojections of the model sets to represent the original Hat and Spectre tilings, which also allows to calculate their Fourier--Bohr coefficients explicitly. Since the windows of the underlying model sets have fractal boundaries, these coefficients need to be computed via an exact renormalisation cocycle in internal space.

math.MG

Broadening Ontologization Design: Embracing Data Pipeline Strategies

Our aim in this paper is to outline how the design space for the ontologization process is broader than current practice would suggest. We point out that engineering processes as well as products need to be designed and identify some components of the design. We investigate the possibility of designing a range of radically new practices implemented as data pipelines, providing examples of the new practices from our work over the last three decades with an outlier methodology, bCLEARer. We also suggest that setting an evolutionary context for ontologization helps one to better understand the nature of these new practices and provides the conceptual scaffolding that shapes fertile processes. Where this evolutionary perspective positions digitalization (the evolutionary emergence of computing technologies) as the latest step in a long evolutionary trail of information transitions. This reframes ontologization as a strategic tool for leveraging the emerging opportunities offered by digitalization.

cs.AI

A classification of intrinsic ergodicity for recognisable random substitution systems

We study a class of dynamical systems generated by random substitutions, which contains both intrinsically ergodic systems and instances with several measures of maximal entropy. In this class, we show that the measures of maximal entropy are classified by invariance under an appropriate symmetry relation. All measures of maximal entropy are fully supported and they are generally not Gibbs measures. We prove that there is a unique measure of maximal entropy if and only if an associated Markov chain is ergodic in inverse time. This Markov chain has finitely many states and all transition matrices are explicitly computable. Thereby, we obtain several sufficient conditions for intrinsic ergodicity that are easy to verify. A practical way to compute the topological entropy in terms of inflation words is extended from previous work to a more general geometric setting.

math.DS

Soundscape Captioning using Sound Affective Quality Network and Large Language Model

We live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring their effects on people, such as the emotions they evoke within a context. To fill this gap, we propose the affective soundscape captioning (ASSC) task, which enables automated soundscape analysis, thus avoiding labour-intensive subjective ratings and surveys in conventional methods. With soundscape captioning, context-aware descriptions are generated for soundscape by capturing the acoustic scenes (ASs), audio events (AEs) information, and the corresponding human affective qualities (AQs). To this end, we propose an automatic soundscape captioner (SoundSCaper) system composed of an acoustic model, i.e. SoundAQnet, and a large language model (LLM). SoundAQnet simultaneously models multi-scale information about ASs, AEs, and perceived AQs, while the LLM describes the soundscape with captions by parsing the information captured with SoundAQnet. SoundSCaper is assessed by two juries of 32 people. In expert evaluation, the average score of SoundSCaper-generated captions is slightly lower than that of two soundscape experts on the evaluation set D1 and the external mixed dataset D2, but not statistically significant. In layperson evaluation, SoundSCaper outperforms soundscape experts in several metrics. In addition to human evaluation, compared to other automated audio captioning systems with and without LLM, SoundSCaper performs better on the ASSC task in several NLP-based metrics. Overall, SoundSCaper performs well in human subjective evaluation and various objective captioning metrics, and the generated captions are comparable to those annotated by soundscape experts.

eess.AS

Rauzy fractals of random substitutions

We develop a theory of Rauzy fractals for random substitutions, which are a generalisation of deterministic substitutions where the substituted image of a letter is determined by a Markov process. We show that a Rauzy fractal can be associated with a given random substitution in a canonical manner, under natural assumptions on the random substitution. Further, we show the existence of a natural measure supported on the Rauzy fractal, which we call the Rauzy measure, that captures geometric and dynamical information. We provide several different constructions for the Rauzy fractal and Rauzy measure, which we show coincide, and ascertain various analytic, dynamical and geometric properties. While the Rauzy fractal is independent of the choice of (non-degenerate) probabilities assigned to a given random substitution, the Rauzy measure captures the explicit choice of probabilities. Moreover, Rauzy measures vary continuously with the choice of probabilities, thus provide a natural means of interpolating between Rauzy fractals of deterministic substitutions. Additionally, we highlight connections between Rauzy fractals and Rauzy measures of random substitutions and related S-adic systems.

math.DS

AI-based soundscape analysis: Jointly identifying sound sources and predicting annoyance

Soundscape studies typically attempt to capture the perception and understanding of sonic environments by surveying users. However, for long-term monitoring or assessing interventions, sound-signal-based approaches are required. To this end, most previous research focused on psycho-acoustic quantities or automatic sound recognition. Few attempts were made to include appraisal (e.g., in circumplex frameworks). This paper proposes an artificial intelligence (AI)-based dual-branch convolutional neural network with cross-attention-based fusion (DCNN-CaF) to analyze automatic soundscape characterization, including sound recognition and appraisal. Using the DeLTA dataset containing human-annotated sound source labels and perceived annoyance, the DCNN-CaF is proposed to perform sound source classification (SSC) and human-perceived annoyance rating prediction (ARP). Experimental findings indicate that (1) the proposed DCNN-CaF using loudness and Mel features outperforms the DCNN-CaF using only one of them. (2) The proposed DCNN-CaF with cross-attention fusion outperforms other typical AI-based models and soundscape-related traditional machine learning methods on the SSC and ARP tasks. (3) Correlation analysis reveals that the relationship between sound sources and annoyance is similar for humans and the proposed AI-based DCNN-CaF model. (4) Generalization tests show that the proposed model's ARP in the presence of model-unknown sound sources is consistent with expert expectations and can explain previous findings from the literature on sound-scape augmentation.

eess.AS

MirrorCalib: Utilizing Human Pose Information for Mirror-based Virtual Camera Calibration

In this paper, we present the novel task of estimating the extrinsic parameters of a virtual camera relative to a real camera in exercise videos with a mirror. This task poses a significant challenge in scenarios where the views from the real and mirrored cameras have no overlap or share salient features. To address this issue, prior knowledge of a human body and 2D joint locations are utilized to estimate the camera extrinsic parameters when a person is in front of a mirror. We devise a modified eight-point algorithm to obtain an initial estimation from 2D joint locations. The 2D joint locations are then refined subject to human body constraints. Finally, a RANSAC algorithm is employed to remove outliers by comparing their epipolar distances to a predetermined threshold. MirrorCalib achieves a rotation error of 1.82{\deg} and a translation error of 69.51 mm on a collected real-world dataset, which outperforms the state-of-art method.

cs.CV

Joint Prediction of Audio Event and Annoyance Rating in an Urban Soundscape by Hierarchical Graph Representation Learning

Sound events in daily life carry rich information about the objective world. The composition of these sounds affects the mood of people in a soundscape. Most previous approaches only focus on classifying and detecting audio events and scenes, but may ignore their perceptual quality that may impact humans' listening mood for the environment, e.g. annoyance. To this end, this paper proposes a novel hierarchical graph representation learning (HGRL) approach which links objective audio events (AE) with subjective annoyance ratings (AR) of the soundscape perceived by humans. The hierarchical graph consists of fine-grained event (fAE) embeddings with single-class event semantics, coarse-grained event (cAE) embeddings with multi-class event semantics, and AR embeddings. Experiments show the proposed HGRL successfully integrates AE with AR for AEC and ARP tasks, while coordinating the relations between cAE and fAE and further aligning the two different grains of AE information with the AR.

eess.AS

On word complexity and topological entropy of random substitution subshifts

We consider word complexity and topological entropy for random substitution subshifts. In contrast to previous work, we do not assume that the underlying random substitution is compatible. We show that the subshift of a primitive random substitution has zero topological entropy if and only if it can be obtained as the subshift of a deterministic substitution, answering in the affirmative an open question of Rust and Spindeler. For constant length primitive random substitutions, we develop a systematic approach to calculating the topological entropy of the associated subshift. Further, we prove lower and upper bounds that hold even without primitivity. For subshifts of non-primitive random substitutions, we show that the complexity function can exhibit features not possible in the deterministic or primitive random setting, such as intermediate growth, and provide a partial classification of the permissible complexity functions for subshifts of constant length random substitutions.

math.DS

Multifractal analysis of measures arising from random substitutions

We study regularity properties of frequency measures arising from random substitutions, which are a generalisation of (deterministic) substitutions where the substituted image of each letter is chosen independently from a fixed finite set. In particular, for a natural class of such measures, we derive a closed-form analytic formula for the $L^q$-spectrum and prove that the multifractal formalism holds. This provides an interesting new class of measures satisfying the multifractal formalism. More generally, we establish results concerning the $L^q$-spectrum of a broad class of frequency measures. We introduce a new notion called the inflation word $L^q$-spectrum of a random substitution and show that this coincides with the $L^q$-spectrum of the corresponding frequency measure for all $q \geq 0$. As an application, we obtain closed-form formulas under separation conditions and recover known results for topological and measure theoretic entropy.

math.DS

Entropy measurement of a strongly coupled quantum dot

The spin 1/2 entropy of electrons trapped in a quantum dot has previously been measured with great accuracy, but the protocol used for that measurement is valid only within a restrictive set of conditions. Here, we demonstrate a novel entropy measurement protocol that is universal for arbitrary mesoscopic circuits and apply this new approach to measure the entropy of a quantum dot hybridized with a reservoir, where Kondo correlations dominate spin physics. The experimental results match closely to numerical renormalization group (NRG) calculations for small and intermediate coupling. For the largest couplings investigated in this work, NRG predicts a suppression of spin entropy at the charge transition due to the formation of a Kondo singlet, but that suppression is not observed in the experiment.

cond-mat.mes-hall

Measure theoretic entropy of random substitution subshifts

Subshifts of deterministic substitutions are ubiquitous objects in dynamical systems and aperiodic order (the mathematical theory of quasicrystals). Two of their most striking features are that they have low complexity (zero topological entropy) and are uniquely ergodic. Random substitutions are a generalisation of deterministic substitutions where the substituted image of a letter is determined by a Markov process. In stark contrast to their deterministic counterparts, subshifts of random substitutions often have positive topological entropy, and support uncountably many ergodic measures. The underlying Markov process singles out one of the ergodic measures, called the frequency measure. Here, we develop new techniques for computing and studying the entropy of these frequency measures. As an application of our results, we obtain closed form formulas for the entropy of frequency measures for a wide range of random substitution subshifts and show that in many cases there exists a frequency measure of maximal entropy. Further, for a class of random substitution subshifts, we prove that this measure is the unique measure of maximal entropy. These subshifts do not satisfy Bowen's specification property or the weaker specification property of Climenhaga and Thompson and hence provide an interesting new class of intrinsically ergodic subshifts.

math.DS