SearcharxivSearch

arXiv subjects

Andrew Jones

Publications and source records attributed to Andrew Jones.

At least 19 recordsLinked to original sources

Enabling tomorrow's planetary defence and space resource economy: Autonomous fleet-based asteroid rendezvous missions

Asteroids preserve the solar system's earliest history and pose real threats to Earth. Strengthening the UK's capabilities to detect, track, and characterise Near-Earth Objects (NEOs) is vital for national security, world-leading planetary science, and future space resource opportunities. The UK has been an influential contributor to planetary defence, from establishing the UK NEO Task Force in 2000 to active roles in the International Asteroid Warning Network (IAWN), the Space Mission Planning Advisory Group (SMPAG), and the development of the National Space Operations Centre (NSpOC). UK scientists contribute to major international asteroid missions including NASA's OSIRIS-REx, DART, Lucy, and Psyche; ESA's Hera and RAMSES; and JAXA's Hayabusa2 and MMX. However, the UK currently lacks dedicated funding streams to deliver asteroid missions. Ground-based observations of asteroids cannot definitively determine the physical characteristics of these objects, which are crucial for impact-risk assessment, deflection strategy, and resource evaluation. This white paper proposes UK leadership in autonomous, low-cost asteroid-rendezvous missions, leveraging technologies developed through the UKRI-funded REMORA programme. We outline four priorities: (1) strengthen NEO detection capabilities; (2) reinforce UK participation in international planetary defence missions; (3) develop autonomous rendezvous and in-situ characterisation technologies; and (4) enable the future space resource economy through targeted asteroid exploration. Together, these actions position the UK to lead rapid, affordable deep-space missions and secure a long-term strategic advantage.

astro-ph.IM

Generalisation and benign over-fitting for linear regression onto random functional covariates

We study theoretical predictive performance of ridge and ridge-less least-squares regression when covariate vectors arise from evaluating $p$ random, means-square continuous functions over a latent metric space at $n$ random and unobserved locations, subject to additive noise. This leads us away from the standard assumption of i.i.d. data to a setting in which the $n$ covariate vectors are exchangeable but not independent in general. Under an assumption of independence across dimensions, $4$-th order moment, and other regularity conditions, we obtain probabilistic bounds on a notion of predictive excess risk adapted to our random functional covariate setting, making use of recent results of Barzilai and Shamir. We derive convergence rates in regimes where $p$ grows suitably fast relative to $n$, illustrating interplay between ingredients of the model in determining convergence behaviour and the role of additive covariate noise in benign-overfitting.

stat.ML

Revealing Flare Energetics and Dynamics with SDO EVE Solar Extreme Ultraviolet Spectral Irradiance Observations

NASA's Solar Dynamics Observatory (SDO) Extreme-ultraviolet Variability Experiment (EVE) has been making solar full-disk extreme ultraviolet (EUV) spectral measurements since 2010 over the spectral range of 6nm to 106nm with 0.1nm spectral resolution and with 10-60sec cadence. A primary motivation for EVE's solar EUV irradiance observations is to provide the important energy input for various studies of Earth's upper atmosphere. For example, the solar EUV creates the ionosphere, heats the thermosphere, and drives photochemistry in Earth's upper atmosphere. In addition, EVE's observations have been a treasure trove for solar EUV flare spectra. While EVE measures the full-disk spectra, the flare spectrum is easily determined as the EVE spectrum minus the pre-flare spectrum, as long as only one flare event is happening at a time. These EVE flare observations provide EUV variability that have been used to study flare phases (including the discovery of the EUV Late Phase flare class), flare energetics (plasma temperature variations), corona heating (plasma abundance changes that support nano-flare heating mechanism), flare dynamics (downwelling and upwelling plasma flows during flares from Doppler shifts), and coronal mass ejections (CME) energetics (CME mass and velocity derived from coronal dimming in some EUV lines). We also introduce a new EVE data product called the EVE Level 4 Lines data product, which provides line profile-fit results for intensity, wavelength shift, and line width for 70 emission features. These emission features are from the chromosphere, transition region, and corona, and so Doppler measurements of those lines can reveal important plasma dynamical behavior during a flare's impulsive phase and gradual phase. With over 10,000 flares detected in the EVE observations, there is still much to study and to learn about solar flare physics using EVE solar EUV spectra.

astro-ph.SR

Unsupervised Attributed Dynamic Network Embedding with Stability Guarantees

Stability for dynamic network embeddings ensures that nodes behaving the same at different times receive the same embedding, allowing comparison of nodes in the network across time. We present attributed unfolded adjacency spectral embedding (AUASE), a stable unsupervised representation learning framework for dynamic networks in which nodes are attributed with time-varying covariate information. To establish stability, we prove uniform convergence to an associated latent position model. We quantify the benefits of our dynamic embedding by comparing with state-of-the-art network representation learning methods on four real attributed networks. To the best of our knowledge, AUASE is the only attributed dynamic embedding that satisfies stability guarantees without the need for ground truth labels, which we demonstrate provides significant improvements for link prediction and node classification.

stat.ML

XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing

HPC and Cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with a shared execution platform that provides transparent access to computing resources, regardless of the underlying cloud or HPC service provider. Bridging HPC and cloud advancements, XaaS presents a unified architecture built on performance-portable containers. Our converged model concentrates on low-overhead, high-performance communication and computing, targeting resource-intensive workloads from climate simulations to machine learning. XaaS lifts the restricted allocation model of Function-as-a-Service (FaaS), allowing users to benefit from the flexibility and efficient resource utilization of serverless while supporting long-running and performance-sensitive workloads from HPC.

cs.DC

Contrastive linear regression

Contrastive dimension reduction methods have been developed for case-control study data to identify variation that is enriched in the foreground (case) data X relative to the background (control) data Y. Here, we develop contrastive regression for the setting when there is a response variable r associated with each foreground observation. This situation occurs frequently when, for example, the unaffected controls do not have a disease grade or intervention dosage but the affected cases have a disease grade or intervention dosage, as in autism severity, solid tumors stages, polyp sizes, or warfarin dosages. Our contrastive regression model captures shared low-dimensional variation between the predictors in the cases and control groups, and then explains the case-specific response variables through the variance that remains in the predictors after shared variation is removed. We show that, in one single-nucleus RNA sequencing dataset on autism severity in postmortem brain samples from donors with and without autism and in another single-cell RNA sequencing dataset on cellular differentiation in chronic rhinosinusitis with and without nasal polyps, our contrastive linear regression performs feature ranking and identifies biologically-informative predictors associated with response that cannot be identified using other approaches

stat.ME

Kernel Density Bayesian Inverse Reinforcement Learning

Inverse reinforcement learning (IRL) methods infer an agent's reward function using demonstrations of expert behavior. A Bayesian IRL approach models a distribution over candidate reward functions, capturing a degree of uncertainty in the inferred reward function. This is critical in some applications, such as those involving clinical data. Typically, Bayesian IRL algorithms require large demonstration datasets, which may not be available in practice. In this work, we incorporate existing domain-specific data to achieve better posterior concentration rates. We study a common setting in clinical and biological applications where we have access to expert demonstrations and known reward functions for a set of training tasks. Our aim is to learn the reward function of a new test task given limited expert demonstrations. Existing Bayesian IRL methods impose restrictions on the form of input data, thus limiting the incorporation of training task data. To better leverage information from training tasks, we introduce kernel density Bayesian inverse reinforcement learning (KD-BIRL). Our approach employs a conditional kernel density estimator, which uses the known reward functions of the training tasks to improve the likelihood estimation across a range of reward functions and demonstration samples. Our empirical results highlight KD-BIRL's faster concentration rate in comparison to baselines, particularly in low test task expert demonstration data regimes. Additionally, we are the first to provide theoretical guarantees of posterior concentration for a Bayesian IRL algorithm. Taken together, this work introduces a principled and theoretically grounded framework that enables Bayesian IRL to be applied across a variety of domains.

cs.LG

Multi-group Gaussian Processes

Gaussian processes (GPs) are pervasive in functional data analysis, machine learning, and spatial statistics for modeling complex dependencies. Modern scientific data sets are typically heterogeneous and often contain multiple known discrete subgroups of samples. For example, in genomics applications samples may be grouped according to tissue type or drug exposure. In the modeling process it is desirable to leverage the similarity among groups while accounting for differences between them. While a substantial literature exists for GPs over Euclidean domains $\mathbb{R}^p$, GPs on domains suitable for multi-group data remain less explored. Here, we develop a multi-group Gaussian process (MGGP), which we define on $\mathbb{R}^p\times \mathscr{C}$, where $\mathscr{C}$ is a finite set representing the group label. We provide general methods to construct valid (positive definite) covariance functions on this domain, and we describe algorithms for inference, estimation, and prediction. We perform simulation experiments and apply MGGP to gene expression data to illustrate the behavior and advantages of the MGGP in the joint modeling of continuous and categorical variables.

stat.ME

Compositional Q-learning for electrolyte repletion with imbalanced patient sub-populations

Reinforcement learning (RL) is an effective framework for solving sequential decision-making tasks. However, applying RL methods in medical care settings is challenging in part due to heterogeneity in treatment response among patients. Some patients can be treated with standard protocols whereas others, such as those with chronic diseases, need personalized treatment planning. Traditional RL methods often fail to account for this heterogeneity, because they assume that all patients respond to the treatment in the same way (i.e., transition dynamics are shared). We introduce Compositional Fitted $Q$-iteration (CFQI), which uses a compositional task structure to represent heterogeneous treatment responses in medical care settings. A compositional task consists of several variations of the same task, each progressing in difficulty; solving simpler variants of the task can enable efficient solving of harder variants. CFQI uses a compositional $Q$-value function with separate modules for each task variant, allowing it to take advantage of shared knowledge while learning distinct policies for each variant. We validate CFQI's performance using a Cartpole environment and use CFQI to recommend electrolyte repletion for patients with and without renal disease. Our results demonstrate that CFQI is robust even in the presence of class imbalance, enabling effective information usage across patient sub-populations. CFQI exhibits great promise for clinical applications in scenarios characterized by known compositional structures.

cs.LG

Spectral embedding for dynamic networks with stability guarantees

We consider the problem of embedding a dynamic network, to obtain time-evolving vector representations of each node, which can then be used to describe changes in behaviour of individual nodes, communities, or the entire graph. Given this open-ended remit, we argue that two types of stability in the spatio-temporal positioning of nodes are desirable: to assign the same position, up to noise, to nodes behaving similarly at a given time (cross-sectional stability) and a constant position, up to noise, to a single node behaving similarly across different times (longitudinal stability). Similarity in behaviour is defined formally using notions of exchangeability under a dynamic latent position network model. By showing how this model can be recast as a multilayer random dot product graph, we demonstrate that unfolded adjacency spectral embedding satisfies both stability conditions. We also show how two alternative methods, omnibus and independent spectral embedding, alternately lack one or the other form of stability.

stat.ML

Full band Monte Carlo simulation of AlInAsSb digital alloys

Avalanche photodiodes fabricated from AlInAsSb grown as a digital alloy exhibit low excess noise. In this paper, we investigate the band structure-related mechanisms that influence impact ionization. Band-structures calculated using an empirical tight-binding method and Monte Carlo simulations reveal that the mini-gaps in the conduction band do not inhibit electron impact ionization. Good agreement between the full band Monte Carlo simulations and measured noise characteristics is demonstrated.

cond-mat.mtrl-sci

Neural Lumigraph Rendering

Novel view synthesis is a challenging and ill-posed inverse rendering problem. Neural rendering techniques have recently achieved photorealistic image quality for this task. State-of-the-art (SOTA) neural volume rendering approaches, however, are slow to train and require minutes of inference (i.e., rendering) time for high image resolutions. We adopt high-capacity neural scene representations with periodic activations for jointly optimizing an implicit surface and a radiance field of a scene supervised exclusively with posed 2D images. Our neural rendering pipeline accelerates SOTA neural volume rendering by about two orders of magnitude and our implicit surface representation is unique in allowing us to export a mesh with view-dependent texture information. Thus, like other implicit surface representations, ours is compatible with traditional graphics pipelines, enabling real-time rendering rates, while achieving unprecedented image quality compared to other surface methods. We assess the quality of our approach using existing datasets as well as high-quality 3D face data captured with a custom multi-camera rig.

cs.CV

Contrastive latent variable modeling with application to case-control sequencing experiments

High-throughput RNA-sequencing (RNA-seq) technologies are powerful tools for understanding cellular state. Often it is of interest to quantify and summarize changes in cell state that occur between experimental or biological conditions. Differential expression is typically assessed using univariate tests to measure gene-wise shifts in expression. However, these methods largely ignore changes in transcriptional correlation. Furthermore, there is a need to identify the low-dimensional structure of the gene expression shift to identify collections of genes that change between conditions. Here, we propose contrastive latent variable models designed for count data to create a richer portrait of differential expression in sequencing data. These models disentangle the sources of transcriptional variation in different conditions, in the context of an explicit model of variation at baseline. Moreover, we develop a model-based hypothesis testing framework that can test for global and gene subset-specific changes in expression. We test our model through extensive simulations and analyses with count-based gene expression data from perturbation and observational sequencing experiments. We find that our methods can effectively summarize and quantify complex transcriptional changes in case-control experimental sequencing data.

stat.ME

Probabilistic Contrastive Principal Component Analysis

Dimension reduction is useful for exploratory data analysis. In many applications, it is of interest to discover variation that is enriched in a "foreground" dataset relative to a "background" dataset. Recently, contrastive principal component analysis (CPCA) was proposed for this setting. However, the lack of a formal probabilistic model makes it difficult to reason about CPCA and to tune its hyperparameter. In this work, we propose probabilistic contrastive principal component analysis (PCPCA), a model-based alternative to CPCA. We discuss how to set the hyperparameter in theory and in practice, and we show several of PCPCA's advantages over CPCA, including greater interpretability, uncertainty quantification and principled inference, robustness to noise and missing data, and the ability to generate data from the model. We demonstrate PCPCA's performance through a series of simulations and case-control experiments with datasets of gene expression, protein expression, and images.

stat.ME

CME Acceleration as a Probe of the Coronal Magnetic Field

By 2050, we expect that CME models will accurately describe, and ideally predict, observed solar eruptions and the propagation of the CMEs through the corona. We describe some of the present known unknowns in observations and models that would need to be addressed in order to reach this goal. We also describe how we might prepare for some of the unknown unknowns that will surely become challenges.

astro-ph.SR

The multilayer random dot product graph

We present a comprehensive extension of the latent position network model known as the random dot product graph to accommodate multiple graphs -- both undirected and directed -- which share a common subset of nodes, and propose a method for jointly embedding the associated adjacency matrices, or submatrices thereof, into a suitable latent space. Theoretical results concerning the asymptotic behaviour of the node representations thus obtained are established, showing that after the application of a linear transformation these converge uniformly in the Euclidean norm to the latent positions with Gaussian error. Within this framework, we present a generalisation of the stochastic block model to a number of different multiple graph settings, and demonstrate the effectiveness of our joint embedding method through several statistical inference tasks in which we achieve comparable or better results than rival spectral methods. Empirical improvements in link prediction over single graph embeddings are exhibited in a cyber-security example.

stat.ML

Spectral embedding of weighted graphs

When analyzing weighted networks using spectral embedding, a judicious transformation of the edge weights may produce better results. To formalize this idea, we consider the asymptotic behavior of spectral embedding for different edge-weight representations, under a generic low rank model. We measure the quality of different embeddings -- which can be on entirely different scales -- by how easy it is to distinguish communities, in an information-theoretic sense. For common types of weighted graphs, such as count networks or p-value networks, we find that transformations such as tempering or thresholding can be highly beneficial, both in theory and in practice.

stat.ML

MinXSS-2 CubeSat mission overview: Improvements from the successful MinXSS-1 mission

The second Miniature X-ray Solar Spectrometer (MinXSS-2) CubeSat, which begins its flight in late 2018, builds on the success of MinXSS-1, which flew from 2016-05-16 to 2017-05-06. The science instrument is more advanced -- now capable of greater dynamic range with higher energy resolution. More data will be captured on the ground than was possible with MinXSS-1 thanks to a sun-synchronous, polar orbit and technical improvements to both the spacecraft and the ground network. Additionally, a new open-source beacon decoder for amateur radio operators is available that can automatically forward any captured MinXSS data to the operations and science team. While MinXSS-1 was only able to downlink about 1 MB of data per day corresponding to a data capture rate of about 1%, MinXSS-2 will increase that by at least a factor of 6. This increase of data capture rate in combination with the mission's longer orbital lifetime will be used to address new science questions focused on how coronal soft X-rays vary over solar cycle timescales and what impact those variations have on the earth's upper atmosphere.

astro-ph.IM