SearcharxivSearch

arXiv subjects

Andrew Wood

Publications and source records attributed to Andrew Wood.

18 recordsLinked to original sources

The Set-Self-Tietze Property

We introduce the set-self-Tietze property, an analogue of the self-Tietze property for upper semi-continuous set-valued functions. A topological space $X$ is self-Tietze, if for every closed $A \subseteq X$ and continuous function $f \colon A \to X$, there is a continuous extension $F \colon X \to X$ of $f$. A topological space $X$ is set-self-Tietze, if for every closed $A \subseteq X$ and upper semi-continuous set-valued function $f \colon A \to 2^X$, there exists an upper semi-continuous set-valued function $F \colon X \to 2^X$ such that $\left. F \right|_A = f$. We show every compact metric space is set-self-Tietze, and that the torus is not self-Tietze.

math.GN

Edge Inversions in $(P_k)$-closed Groups

We construct $(P_2)$-closed groups acting on $T_3$ in which all edge inversions have infinite order. This provides a negative answer to a question posed by Tornier. We also construct a family of $(P_2)$-closed groups for which the smallest order of an edge inversion is an arbitrarily high finite number.

math.GR

Transitivity in CR-Dynamical Systems

A CR-dynamical system is a pair $(X, G)$, where $X$ is a compact metric space and $G$ is a closed relation (CR) on $X$. In this paper, we introduce a new type of transitive point and transitivity in CR-dynamical systems. We develop a new tool called transitivity trees, which we use to determine the relationship between the different types of transitive points.

math.DS

Shadowing in CR-Dynamical Systems

A CR-dynamical system is a pair $(X, G)$, where $X$ is a non-empty compact Hausdorff space with uniformity $\mathscr{U}$ and $G$ is a closed relation on $X$. In this paper we introduce the $(i, j)$-shadowing properties in CR-dynamical systems, which generalises the shadowing property from topological dynamical systems $(X, f)$. This extends previous work on shadowing in set-valued dynamical systems.

math.DS

PhenoAssistant: A Conversational Multi-Agent AI System for Automated Plant Phenotyping

Plant phenotyping increasingly relies on (semi-)automated image-based analysis workflows to improve its accuracy and scalability. However, many existing solutions remain overly complex, difficult to reimplement and maintain, and pose high barriers for users without substantial computational expertise. To address these challenges, we introduce PhenoAssistant: a pioneering AI-driven system that streamlines plant phenotyping via intuitive natural language interaction. PhenoAssistant leverages a large language model to orchestrate a curated toolkit supporting tasks including automated phenotype extraction, data visualisation and automated model training. We validate PhenoAssistant through several representative case studies and a set of evaluation tasks. By significantly lowering technical hurdles, PhenoAssistant underscores the promise of AI-driven methodologies to democratising AI adoption in plant biology.

cs.MA

ZipNN: Lossless Compression for AI Models

With the growth of model sizes and the scale of their deployment, their sheer size burdens the infrastructure requiring more network and more storage to accommodate these. While there is a vast model compression literature deleting parts of the model weights for faster inference, we investigate a more traditional type of compression - one that represents the model in a compact form and is coupled with a decompression algorithm that returns it to its original form and size - namely lossless compression. We present ZipNN a lossless compression tailored to neural networks. Somewhat surprisingly, we show that specific lossless compression can gain significant network and storage reduction on popular models, often saving 33% and at times reducing over 50% of the model size. We investigate the source of model compressibility and introduce specialized compression variants tailored for models that further increase the effectiveness of compression. On popular models (e.g. Llama 3) ZipNN shows space savings that are over 17% better than vanilla compression while also improving compression and decompression speeds by 62%. We estimate that these methods could save over an ExaByte per month of network traffic downloaded from a large model hub like Hugging Face.

cs.LG

AI Age Discrepancy: A Novel Parameter for Frailty Assessment in Kidney Tumor Patients

Kidney cancer is a global health concern, and accurate assessment of patient frailty is crucial for optimizing surgical outcomes. This paper introduces AI Age Discrepancy, a novel metric derived from machine learning analysis of preoperative abdominal CT scans, as a potential indicator of frailty and postoperative risk in kidney cancer patients. This retrospective study of 599 patients from the 2023 Kidney Tumor Segmentation (KiTS) challenge dataset found that a higher AI Age Discrepancy is significantly associated with longer hospital stays and lower overall survival rates, independent of established factors. This suggests that AI Age Discrepancy may provide valuable insights into patient frailty and could thus inform clinical decision-making in kidney cancer treatment.

cs.CV

Lossless and Near-Lossless Compression for Foundation Models

With the growth of model sizes and scale of their deployment, their sheer size burdens the infrastructure requiring more network and more storage to accommodate these. While there is a vast literature about reducing model sizes, we investigate a more traditional type of compression -- one that compresses the model to a smaller form and is coupled with a decompression algorithm that returns it to its original size -- namely lossless compression. Somewhat surprisingly, we show that such lossless compression can gain significant network and storage reduction on popular models, at times reducing over $50\%$ of the model size. We investigate the source of model compressibility, introduce compression variants tailored for models and categorize models to compressibility groups. We also introduce a tunable lossy compression technique that can further reduce size even on the less compressible models with little to no effect on the model accuracy. We estimate that these methods could save over an ExaByte per month of network traffic downloaded from a large model hub like HuggingFace.

cs.LG

Improved prediction of hiking speeds using a data driven approach

Hikers and hillwalkers typically use the gradient in the direction of travel (walking slope) as the main variable in established methods for predicting walking time (via the walking speed) along a route. Research into fell-running has suggested further variables which may improve speed algorithms in this context; the gradient of the terrain (hill slope) and the level of terrain obstruction. Recent improvements in data availability, as well as widespread use of GPS tracking now make it possible to explore these variables in a walking speed model at a sufficient scale to test statistical significance. We tested various established models used to predict walking speed against public GPS data from almost 88,000 km of UK walking / hiking tracks. Tracks were filtered to remove breaks and non-walking sections. A new generalised linear model (GLM) was then used to predict walking speeds. Key differences between the GLM and established rules were that the GLM considered the gradient of the terrain (hill slope) irrespective of walking slope, as well as the terrain type and level of terrain obstruction in off-road travel. All of these factors were shown to be highly significant, and this is supported by a lower root-mean-square-error compared to existing functions. We also observed an increase in RMSE between the GLM and established methods as hill slope increases, further supporting the importance of this variable.

stat.AP

Non-Volatile Memory Accelerated Posterior Estimation

Bayesian inference allows machine learning models to express uncertainty. Current machine learning models use only a single learnable parameter combination when making predictions, and as a result are highly overconfident when their predictions are wrong. To use more learnable parameter combinations efficiently, these samples must be drawn from the posterior distribution. Unfortunately computing the posterior directly is infeasible, so often researchers approximate it with a well known distribution such as a Gaussian. In this paper, we show that through the use of high-capacity persistent storage, models whose posterior distribution was too big to approximate are now feasible, leading to improved predictions in downstream tasks.

cs.LG

Non-Volatile Memory Accelerated Geometric Multi-Scale Resolution Analysis

Dimensionality reduction algorithms are standard tools in a researcher's toolbox. Dimensionality reduction algorithms are frequently used to augment downstream tasks such as machine learning, data science, and also are exploratory methods for understanding complex phenomena. For instance, dimensionality reduction is commonly used in Biology as well as Neuroscience to understand data collected from biological subjects. However, dimensionality reduction techniques are limited by the von-Neumann architectures that they execute on. Specifically, data intensive algorithms such as dimensionality reduction techniques often require fast, high capacity, persistent memory which historically hardware has been unable to provide at the same time. In this paper, we present a re-implementation of an existing dimensionality reduction technique called Geometric Multi-Scale Resolution Analysis (GMRA) which has been accelerated via novel persistent memory technology called Memory Centric Active Storage (MCAS). Our implementation uses a specialized version of MCAS called PyMM that provides native support for Python datatypes including NumPy arrays and PyTorch tensors. We compare our PyMM implementation against a DRAM implementation, and show that when data fits in DRAM, PyMM offers competitive runtimes. When data does not fit in DRAM, our PyMM implementation is still able to process the data.

cs.LG

What is Learned in Knowledge Graph Embeddings?

A knowledge graph (KG) is a data structure which represents entities and relations as the vertices and edges of a directed graph with edge types. KGs are an important primitive in modern machine learning and artificial intelligence. Embedding-based models, such as the seminal TransE [Bordes et al., 2013] and the recent PairRE [Chao et al., 2020] are among the most popular and successful approaches for representing KGs and inferring missing edges (link completion). Their relative success is often credited in the literature to their ability to learn logical rules between the relations. In this work, we investigate whether learning rules between relations is indeed what drives the performance of embedding-based methods. We define motif learning and two alternative mechanisms, network learning (based only on the connectivity of the KG, ignoring the relation types), and unstructured statistical learning (ignoring the connectivity of the graph). Using experiments on synthetic KGs, we show that KG models can learn motifs and how this ability is degraded by non-motif (noise) edges. We propose tests to distinguish the contributions of the three mechanisms to performance, and apply them to popular KG benchmarks. We also discuss an issue with the standard performance testing protocol and suggest an improvement. To appear in the proceedings of Complex Networks 2021.

cs.AI

Human Exposure to Radiofrequency Energy above 6 GHz: Review of Computational Dosimetry Studies

International guidelines/standards for human protection from electromagnetic fields have been revised recently, especially for frequencies above 6 GHz where new wireless communication systems have been deployed. Above this frequency a new physical quantity "absorbed/epithelia power density" has been adopted as a dose metric. Then, the permissible level of external field strength/power density is derived for practical assessment. In addition, a new physical quantity, fluence or absorbed energy density, is introduced for protection from brief pulses (especially for shorter than 10 sec). These limits were explicitly designed to avoid excessive increases in tissue temperature, based on electromagnetic and thermal modeling studies but supported by experimental data where available. This paper reviews the studies on the computational modeling/dosimetry which are related to the revision of the guidelines/standards. The comparisons with experimental data as well as an analytic solution are also been presented. Future research needs and additional comments on the revision will also be mentioned.

physics.med-ph

GymFG: A Framework with a Gym Interface for FlightGear

Over the past decades, progress in deployable autonomous flight systems has slowly stagnated. This is reflected in today's production air-crafts, where pilots only enable simple physics-based systems such as autopilot for takeoff, landing, navigation, and terrain/traffic avoidance. Evidently, autonomy has not gained the trust of the community where higher problem complexity and cognitive workload are required. To address trust, we must revisit the process for developing autonomous capabilities: modeling and simulation. Given the prohibitive costs for live tests, we need to prototype and evaluate autonomous aerial agents in a high fidelity flight simulator with autonomous learning capabilities applicable to flight systems: such a open-source development platform is not available. As a result, we have developed GymFG: GymFG couples and extends a high fidelity, open-source flight simulator and a robust agent learning framework to facilitate learning of more complex tasks. Furthermore, we have demonstrated the use of GymFG to train an autonomous aerial agent using Imitation Learning. With GymFG, we can now deploy innovative ideas to address complex problems and build the trust necessary to move prototypes to the real-world.

cs.AI

Detecting Speech Act Types in Developer Question/Answer Conversations During Bug Repair

This paper targets the problem of speech act detection in conversations about bug repair. We conduct a "Wizard of Oz" experiment with 30 professional programmers, in which the programmers fix bugs for two hours, and use a simulated virtual assistant for help. Then, we use an open coding manual annotation procedure to identify the speech act types in the conversations. Finally, we train and evaluate a supervised learning algorithm to automatically detect the speech act types in the conversations. In 30 two-hour conversations, we made 2459 annotations and uncovered 26 speech act types. Our automated detection achieved 69% precision and 50% recall. The key application of this work is to advance the state of the art for virtual assistants in software engineering. Virtual assistant technology is growing rapidly, though applications in software engineering are behind those in other areas, largely due to a lack of relevant data and experiments. This paper targets this problem in the area of developer Q/A conversations about bug repair.

cs.SE

Quantum-correlated photons from semiconductor cavity polaritons

Over the past decade, exciton-polaritons in semiconductor microcavities have attracted a great deal of interest as a driven-dissipative quantum fluid. These systems offer themselves as a versatile platform for performing Hamiltonian simulations with light, as well as for experimentally realizing nontrivial out-of-equilibrium phase transitions. In addition, polaritons exhibit a sizeable mutual interaction strength that opens up a whole range of possibilities in the context of quantum state generation. While squeezed light emission from polaritons has been reported previously, the granular nature of polaritons has not been observed to date. The latter capability is particularly attractive for realizing strongly correlated many-body quantum states of light on scalable arrays of coupled cavities. Here we demonstrate that by optically confining polaritons to a very small effective mode volume, one can reach the weak blockade regime, in which the nonlinearity turns strong enough to become significant at the few particle level, and thus produce a non-negligible antibunching in the emitted photons statistics. Our results act as a door opener for accessing the newly emerging field of quantum polaritonics, and as a proof of principle that optically confined exciton-polaritons can be considered as a realistic, new strategy to generate single photons.

cond-mat.mes-hall

Refutation of Aslam's Proof that NP = P

Aslam presents an algorithm he claims will count the number of perfect matchings in any incomplete bipartite graph with an algorithm in the function-computing version of NC, which is itself a subset of FP. Counting perfect matchings is known to be #P-complete; therefore if Aslam's algorithm is correct, then NP=P. However, we show that Aslam's algorithm does not correctly count the number of perfect matchings and offer an incomplete bipartite graph as a concrete counter-example.

cs.CC

Effect of door delay on aircraft evacuation time

The recent commercial launch of twin-deck Very Large Transport Aircraft (VLTA) such as the Airbus A380 has raised questions concerning the speed at which they may be evacuated. The abnormal height of emergency exits on the upper deck has led to speculation that emotional factors such as fear may lead to door delay, and thus play a significant role in increasing overall evacuation time. Full-scale evacuation tests are financially expensive and potentially hazardous, and systematic studies of the evacuation of VLTA are rare. Here we present a computationally cheap agent-based framework for the general simulation of aircraft evacuation, and apply it to the particular case of the Airbus A380. In particular, we investigate the effect of door delay, and conclude that even a moderate average delay can lead to evacuation times that exceed the maximum for safety certification. The model suggests practical ways to minimise evacuation time, as well as providing a general framework for the simulation of evacuation.

cs.MA