SearcharxivSearch

arXiv subjects

John Brennan

Publications and source records attributed to John Brennan.

At least 19 recordsLinked to original sources

Tabular foundation models for non-tabular tasks

Tabular foundation models (TFMs) have recently emerged as a promising paradigm for machine learning on tabular data, offering the ability to generalize across datasets without task-specific training. Since many machine learning datasets can be represented as tables, this raises the question: does TFM capability extend beyond tasks traditionally regarded as tabular? We address this question by using TabPFN v3 on three non-tabular classification problems: handwritten digit recognition on MNIST, language identification of French and German words, and image classification on Tiny ImageNet. In each case, the original data are represented as rows of a table and classification is formulated as prediction of a missing label. We evaluate performance as a function of the number of context samples provided to the pretrained model, with no additional training or fine-tuning. Despite having no explicit access to the spatial or sequential structure characterizing the data, TabPFN v3 in some cases achieves accuracies comparable with that of models or methods geared specifically toward the corresponding tasks.

cs.LG

SEEDZ: Rapid Galaxy Assembly as a Pathway to Supermassive Stars, Dense Stellar Environments and Massive Black Hole Seeds

We investigate the assembly history of early galaxies in the SEEDZ hydrodynamic simulations, to investigate the high inflow rates believed to be required for the formation of supermassive stars (SMSs), dense stellar clusters and subsequently heavy seed black holes. Using a heavy seed formation criteria of $>$1 M$_\odot$ yr$^{-1}$ flowing into 10 pc regions, we find that heavy seeds form in halos that grow rapidly compared to those halos that never meet the criteria. Halos with growth rates of $\gtrsim$1 M$_\odot$ yr$^{-1}$ at their virial radius (scales of a few hundred pc) are able to sustain a flow rate of 0.1 M$_\odot$ yr$^{-1}$ into the inner 1 pc of the halo, maintaining higher density environments within the central 10 - 100~pc. These halos continue to grow rapidly after their initial collapse, typically forming heavy seeds $\sim$100 Myr after forming their first stars and stellar mass black holes. By $z=10$, most heavy seeds form in regions of near-solar metallicity, although a minority of heavy seeds do continue to form in low metallicity (10$^{-2}$ Z$_\odot$) regions. Under the assumption that a SMS forms as the progenitor to a heavy seed if it forms in a region of low (10$^{-2}$ Z$_\odot$) metallicity, and can sustain high accretion rates above 0.02 M$_\odot$ yr$^{-1}$ throughout the SMS lifetime of 2 Myr, we find a number density of SMSs of 0.1 cMpc$^{-3}$, meaning that only a fraction of 10$^{-4}$ of these SMSs would need to be visible to JWST to account for the observed population of Little Red Dot galaxies.

astro-ph.GA

Black Hole Feedback, Galaxy Quenching and Outflows at Cosmic Dawn: Analysis of the SEEDZ Simulations

Here we analyse the growth and feedback effects of massive black holes (MBHs) in the SEEDZ simulations. The most massive black holes grow to masses of $\sim10^{6}$ M$_\odot$ by $z=12.5$ during short bursts of super-Eddington accretion, sustained over a period of 5-30 Myr. We find that the determining factor that cuts off this initial growth is feedback from the MBH itself, rather than nearby supernovae or exhausting the available gas reservoir. Our simulations show that for the most actively accreting MBHs, feedback completely evacuates the gas from the host halo and ejects it into the inter-galactic medium. Despite implementing a relatively weak feedback model, the energy injected into the gas surrounding the MBH exceeds the binding energy of the halo. These results either indicate that MBH feedback in the early ($\Lambda$CDM) Universe is much weaker than previously assumed, or that at least some of the high redshift galaxies we currently observe with JWST formed via a two-step process, whereby a MBH initially quenches its host galaxy and later reconstitutes its baryonic reservoir, either through mergers with gas rich galaxies or from accretion from the cosmic web. Moreover, the maximum black hole masses that emerge in SEEDZ are effectively set by a combination of MBH feedback modelling and the binding potential of the host halo. Unless feedback is extremely ineffective at early times (for example if growth is merger dominated rather than accretion dominated or feedback is contained close to the MBH) then the maximum mass of black holes at redshift before 12.5 should not significantly exceed $10^6$ M$_\odot$.

astro-ph.GA

Halo mass functions at high redshift

Recent JWST observations of very early galaxies, at $\rm{z \gtrsim 10}$, have led to claims that tension exists between the sizes and luminosities of high-redshift galaxies and what is predicted by standard $Λ$CDM models. Here we use the adaptive mesh refinement code $\texttt{Enzo}$ and the N-body smoothed particle hydrodynamics code $\texttt{SWIFT}$ to compare (semi-)analytic halo mass functions against the results of direct N-body models at high redshift. In particular, our goal is to investigate the variance between standard halo mass functions derived from (semi-)analytic formulations and N-body calculations and to determine what role any discrepancy may play in driving tensions between observations and theory. We find that the difference between direct N-body calculations and (semi-) analytic halo mass function fits is less than a factor of 2 (at $\rm{z \sim 10}$) within the mass range of galaxies currently being observed by JWST, and is therefore not a dominant source of error when comparing theory and observation at high redshift.

astro-ph.CO

The SEEDZ Simulations: Methodology and First Results on Massive Black Hole Seeding and Early Galaxy Growth

Here we introduce the SEEDZ simulations, a suite of cosmological hydrodynamic simulations exploring the formation and growth of the first massive black holes in the Universe. SEEDZ includes models for Population III star formation, supernovae explosions and the resulting formation of light seed black holes, metal enrichment and subsequent Population II star formation, heavy seed black hole formation, Eddington and super-Eddington accretion schemes as well as black hole feedback. In this paper, we cover the overall methodologies employed and present our current results at $z=15$. Our main result so far is that black holes initially grow faster than their host galaxy, and hence over-massive black holes are a feature of the high-redshift Universe. The fundamental black hole-galaxy relationships we observe at $z = 0$ (especially the M$_{\rm BH}$ - M$_*$ relationship) likely only emerge in more mature galaxies. At high-redshift, that relationship has not yet been established. We find that even at these high redshifts, MBHs can grow from their initial heavy seed mass of $\sim$10$^4$ M$_\odot$ up to 10$^6$ M$_\odot$. At the high end of our MBH masses, our simulated galaxy M$_{\rm BH}$ - M$_*$ relations match the observed high redshift trends i.e. over-massive BHs with M$_{\rm BH}$/M$_{\rm star} \sim 10^{-2}$. This initial set of simulations will continue to run down to $z=10$, where we will perform a comprehensive comparison of simulated MBH number densities and M$_{\rm BH}$ - M$_*$ relations with JWST observations. Further simulations with higher resolution will then follow.

astro-ph.GA

Bridging Machine Learning and Cosmological Simulations: Using Neural Operators to emulate Chemical Evolution

The computational expense of solving non-equilibrium chemistry equations in astrophysical simulations poses a significant challenge, particularly in high-resolution, large-scale cosmological models. In this work, we explore the potential of machine learning, specifically Neural Operators, to emulate the Grackle chemistry solver, which is widely used in cosmological hydrodynamical simulations. Neural Operators offer a mesh-free, data-driven approach to approximate solutions to coupled ordinary differential equations governing chemical evolution, gas cooling, and heating. We construct and train multiple Neural Operator architectures (DeepONet variants) using a dataset derived from cosmological simulations to optimize accuracy and efficiency. Our results demonstrate that the trained models accurately reproduce Grackle's outputs with an average error of less than 0.6 dex in most cases, though deviations increase in highly dynamic chemical environments. Compared to Grackle, the machine learning models provide computational speedups of up to a factor of six in large-scale simulations, highlighting their potential for reducing computational bottlenecks in astrophysical modeling. However, challenges remain, particularly in iterative applications where accumulated errors can lead to numerical instability. Additionally, the performance of these machine learning models is constrained by their need for well-represented training datasets and the limited extrapolation capabilities of deep learning methods. While promising, further development is required for Neural Operator-based emulators to be fully integrated into astrophysical simulations. Future work should focus on improving stability over iterative timesteps and optimizing implementations for hardware acceleration. This study provides an initial step toward the broader adoption of machine learning approaches in astrophysical chemistry solvers.

astro-ph.IM

Predicting the number density of heavy seed massive black holes due to an intense Lyman-Werner field

The recent detections of a large number of candidate active galactic nuclei at high redshift (i.e. $z \gtrsim 4$) has increased speculation that heavy seed massive black hole formation may be a required pathway. Here we re-implement the so-called Lyman-Werner (LW) channel model of Dijkstra et al. (2014) to calculate the expected number density of massive black holes formed through this channel. We further enhance this model by extracting information relevant to the model from the $\texttt{Renaissance}$ simulation suite. $\texttt{Renaissance}$ is a high-resolution suite of simulations ideally positioned to probe the high-$z$ Universe. Finally, we compare the LW-only channel against other models in the literature. We find that the LW-only channel results in a peak number density of massive black holes of approximately $\rm{10^{-4} \ cMpc^{-3}}$ at $z \sim 10$. Given the growth requirements and the duty cycle of active galactic nuclei, this means that the LW-only is likely incompatible with recent JWST measurements and can, at most, be responsible for only a small subset of high-$z$ active galactic nuclei. Other models from the literature (e.g. rapid assembly; relative velocities between baryons and dark matter) seem therefore better positioned, at present, to explain the high frequency of massive black holes at high $z$.

astro-ph.CO

On the Use of WGANs for Super-Resolution in Dark-Matter Simulations

Super-resolution techniques have the potential to reduce the computational cost of cosmological and astrophysical simulations. This can be achieved by enabling traditional simulation methods to run at lower resolution and then efficiently computing high-resolution data corresponding to the simulated low-resolution data. In this work, we investigate the application of a Wasserstein Generative Adversarial Network (WGAN) model, previously proposed in the literature, to increase the particle resolution of dark-matter-only simulations. We reproduce prior results, showing the WGAN model successfully generates high-resolution data with summary statistics, including the power spectrum and halo mass function, that closely match those of true high-resolution simulations. However, we also identify a limitation of the WGAN model in the form of smeared features in generated high-resolution data, particularly in the shapes of dark-matter halos and filaments. This limitation points to a potential weakness of the proposed WGAN-based super-resolution method in capturing the detailed structure of halos, and underscores the need for further development in applying such models to cosmological data.

astro-ph.GA

On resolving meso-scale calculations of pore-collapse-generated hotspots in energetic crystals for consistency with atomistic models

Meso-scale calculations of pore collapse and hotspot formation in energetic crystals provide closure models to macro-scale hydrocodes for predicting the shock sensitivity of energetic materials. To this end, previous works obtained atomistics-consistent material models for two common energetic crystals, HMX and RDX, such that pore collapse calculations adhered closely to molecular dynamics (MD) results on key features of energy localization, particularly the profiles of the collapsing pores, appearance of shear bands, and the transition from viscoplastic to hydrodynamic collapse. However, some important aspects such as the temperature distributions in the hotspot were not as well captured. One potential issue was noted but not resolved adequately in those works, namely the grid resolution that should be employed in the meso-scale calculations for various pore sizes and shock strengths. Conventional computational mechanics guidelines for selecting meshes as fine as possible, balancing computational effort, accuracy and grid independence, were shown not to produce physically consistent features associated with shear localization. Here, we examine the physics of pore collapse, shear band evolution and structure, and hotspot formation, for both HMX and RDX; we then evaluate under what conditions atomistics-consistent models yield physically correct (considering MD as ground truth) hotspots for a range of pore diameters, from nm to microns, and for a wide range of shock strengths. The study provides insights into the effects of pore size and shock strength on pore collapse and hotspots, identifying aspects such as size-independent behaviors, and proportion of energy contained in shear as opposed to jet impact-heated regions of the hotspot. Areas for further improvement of atomistics-consistent material models are also indicated.

cond-mat.mes-hall

The Kitaev honeycomb model on surfaces of genus $g \geq 2$

We present a construction of the Kitaev honeycomb lattice model on an arbitrary higher genus surface. We first generalize the exact solution of the model based on the Jordan-Wigner fermionization to a surface with genus $g = 2$, and then use this as a basic module to extend the solution to lattices of arbitrary genus. We demonstrate our method by calculating the ground states of the model in both the Abelian doubled $\mathbb{Z}_2$ phase and the non-Abelian Ising topological phase on lattices with the genus up to $g = 6$. We verify the expected ground state degeneracy of the system in both topological phases and further illuminate the role of fermionic parity in the Abelian phase.

quant-ph

On the Complexity of Object Detection on Real-world Public Transportation Images for Social Distancing Measurement

Social distancing in public spaces has become an essential aspect in helping to reduce the impact of the COVID-19 pandemic. Exploiting recent advances in machine learning, there have been many studies in the literature implementing social distancing via object detection through the use of surveillance cameras in public spaces. However, to date, there has been no study of social distance measurement on public transport. The public transport setting has some unique challenges, including some low-resolution images and camera locations that can lead to the partial occlusion of passengers, which make it challenging to perform accurate detection. Thus, in this paper, we investigate the challenges of performing accurate social distance measurement on public transportation. We benchmark several state-of-the-art object detection algorithms using real-world footage taken from the London Underground and bus network. The work highlights the complexity of performing social distancing measurement on images from current public transportation onboard cameras. Further, exploiting domain knowledge of expected passenger behaviour, we attempt to improve the quality of the detections using various strategies and show improvement over using vanilla object detection alone.

cs.CV

Tensor Network Circuit Simulation at Exascale

Tensor network methods are incredibly effective for simulating quantum circuits. This is due to their ability to efficiently represent and manipulate the wave-functions of large interacting quantum systems. We describe the challenges faced when scaling tensor network simulation approaches to Exascale compute platforms and introduce QuantEx, a framework for tensor network circuit simulation at Exascale.

quant-ph

Deploying Containerized QuantEx Quantum Simulation Software on HPC Systems

The simulation of quantum circuits using the tensor network method is very computationally demanding and requires significant High Performance Computing (HPC) resources to find an efficient contraction order and to perform the contraction of the large tensor networks. In addition, the researchers want a workflow that is easy to customize, reproduce and migrate to different HPC systems. In this paper, we discuss the issues associated with the deployment of the QuantEX quantum computing simulation software within containers on different HPC systems. Also, we compare the performance of the containerized software with the software running on bare metal.

cs.DC

Not Half Bad: Exploring Half-Precision in Graph Convolutional Neural Networks

With the growing significance of graphs as an effective representation of data in numerous applications, efficient graph analysis using modern machine learning is receiving a growing level of attention. Deep learning approaches often operate over the entire adjacency matrix -- as the input and intermediate network layers are all designed in proportion to the size of the adjacency matrix -- leading to intensive computation and large memory requirements as the graph size increases. It is therefore desirable to identify efficient measures to reduce both run-time and memory requirements allowing for the analysis of the largest graphs possible. The use of reduced precision operations within the forward and backward passes of a deep neural network along with novel specialised hardware in modern GPUs can offer promising avenues towards efficiency. In this paper, we provide an in-depth exploration of the use of reduced-precision operations, easily integrable into the highly popular PyTorch framework, and an analysis of the effects of Tensor Cores on graph convolutional neural networks. We perform an extensive experimental evaluation of three GPU architectures and two widely-used graph analysis tasks (vertex classification and link prediction) using well-known benchmark and synthetically generated datasets. Thus allowing us to make important observations on the effects of reduced-precision operations and Tensor Cores on computational and memory usage of graph convolutional neural networks -- often neglected in the literature.

cs.LG

Gradient descent with momentum --- to accelerate or to super-accelerate?

We consider gradient descent with `momentum', a widely used method for loss function minimization in machine learning. This method is often used with `Nesterov acceleration', meaning that the gradient is evaluated not at the current position in parameter space, but at the estimated position after one step. In this work, we show that the algorithm can be improved by extending this `acceleration' --- by using the gradient at an estimated position several steps ahead rather than just one step ahead. How far one looks ahead in this `super-acceleration' algorithm is determined by a new hyperparameter. Considering a one-parameter quadratic loss function, the optimal value of the super-acceleration can be exactly calculated and analytically estimated. We show explicitly that super-accelerating the momentum algorithm is beneficial, not only for this idealized problem, but also for several synthetic loss landscapes and for the MNIST classification task with neural networks. Super-acceleration is also easy to incorporate into adaptive algorithms like RMSProp or Adam, and is shown to improve these algorithms.

cs.LG

Temporal Neighbourhood Aggregation: Predicting Future Links in Temporal Graphs via Recurrent Variational Graph Convolutions

Graphs have become a crucial way to represent large, complex and often temporal datasets across a wide range of scientific disciplines. However, when graphs are used as input to machine learning models, this rich temporal information is frequently disregarded during the learning process, resulting in suboptimal performance on certain temporal infernce tasks. To combat this, we introduce Temporal Neighbourhood Aggregation (TNA), a novel vertex representation model architecture designed to capture both topological and temporal information to directly predict future graph states. Our model exploits hierarchical recurrence at different depths within the graph to enable exploration of changes in temporal neighbourhoods, whilst requiring no additional features or labels to be present. The final vertex representations are created using variational sampling and are optimised to directly predict the next graph in the sequence. Our claims are reinforced by extensive experimental evaluation on both real and synthetic benchmark datasets, where our approach demonstrates superior performance compared to competing methods, out-performing them at predicting new temporal edges by as much as 23% on real-world datasets, whilst also requiring fewer overall model parameters.

cs.SI

Predicting the Computational Cost of Deep Learning Models

Deep learning is rapidly becoming a go-to tool for many artificial intelligence problems due to its ability to outperform other approaches and even humans at many problems. Despite its popularity we are still unable to accurately predict the time it will take to train a deep learning network to solve a given problem. This training time can be seen as the product of the training time per epoch and the number of epochs which need to be performed to reach the desired level of accuracy. Some work has been carried out to predict the training time for an epoch -- most have been based around the assumption that the training time is linearly related to the number of floating point operations required. However, this relationship is not true and becomes exacerbated in cases where other activities start to dominate the execution time. Such as the time to load data from memory or loss of performance due to non-optimal parallel execution. In this work we propose an alternative approach in which we train a deep learning network to predict the execution time for parts of a deep learning network. Timings for these individual parts can then be combined to provide a prediction for the whole execution time. This has advantages over linear approaches as it can model more complex scenarios. But, also, it has the ability to predict execution times for scenarios unseen in the training data. Therefore, our approach can be used not only to infer the execution time for a batch, or entire epoch, but it can also support making a well-informed choice for the appropriate hardware and model.

cs.LG

Temporal Graph Offset Reconstruction: Towards Temporally Robust Graph Representation Learning

Graphs are a commonly used construct for representing relationships between elements in complex high dimensional datasets. Many real-world phenomenon are dynamic in nature, meaning that any graph used to represent them is inherently temporal. However, many of the machine learning models designed to capture knowledge about the structure of these graphs ignore this rich temporal information when creating representations of the graph. This results in models which do not perform well when used to make predictions about the future state of the graph -- especially when the delta between time stamps is not small. In this work, we explore a novel training procedure and an associated unsupervised model which creates graph representations optimised to predict the future state of the graph. We make use of graph convolutional neural networks to encode the graph into a latent representation, which we then use to train our temporal offset reconstruction method, inspired by auto-encoders, to predict a later time point -- multiple time steps into the future. Using our method, we demonstrate superior performance for the task of future link prediction compared with none-temporal state-of-the-art baselines. We show our approach to be capable of outperforming non-temporal baselines by 38% on a real world dataset.

cs.SI