SearcharxivSearch

arXiv subjects

Jonathan Coles

Publications and source records attributed to Jonathan Coles.

10 recordsLinked to original sources

An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience

Large Language Models (LLMs) have surged as a transformative technology for science and society, prompting governments worldwide to pursue sovereign AI capabilities that ensure data compliance and cultural representation. However, the associated capital costs and engineering complexity required to train these models have largely restricted such capabilities to the private sector, leaving a significant gap for public institutions. This paper details the engineering journey behind training Apertus, a fully open multilingual foundation model, on the Alps supercomputer. Representing a first-of-its-kind achievement for academia at the 70B parameter scale, we successfully deployed a massive pre-training campaign on one of Europe's largest systems for open science, powered by NVIDIA GH200 Grace Hopper Superchips. We detail the challenges encountered in readying HPC infrastructure for training AI models, from overcoming storage bottlenecks to stabilizing large-scale interconnects, and the lessons learned in transforming a supercomputer into a resilient software-defined Machine Learning Platform. Finally, we discuss the post-training requirements and evolution of our Machine Learning platform, outlining how this initial release lays the groundwork for a sustained, iterative operational capability, in particular for fine tuning foundation models, that extends well beyond a single model training run.

cs.DC

Computing the Full Earth System at 1 km Resolution

We present the first-ever global simulation of the full Earth system at 1.25 km grid spacing, achieving highest time compression with an unseen number of degrees of freedom. Our model captures the flow of energy, water, and carbon through key components of the Earth system: atmosphere, ocean, and land. To achieve this landmark simulation, we harness the power of 8192 GPUs on Alps and 20480 GPUs on JUPITER, two of the world's largest GH200 superchip installations. We use both the Grace CPUs and Hopper GPUs by carefully balancing Earth's components in a heterogeneous setup and optimizing acceleration techniques available in ICON's codebase. We show how separation of concerns can reduce the code complexity by half while increasing performance and portability. Our achieved time compression of 145.7 simulated days per day enables long studies including full interactions in the Earth system and even outperforms earlier atmosphere-only simulations at a similar resolution.

physics.ao-ph

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingual representation. Unlike many prior models that release weights without reproducible data pipelines or regard for content-owner rights, Apertus models are pretrained exclusively on openly available data, retroactively respecting `robots.txt` exclusions and filtering for non-permissive, toxic, and personally identifiable content. To mitigate risks of memorization, we adopt the Goldfish objective during pretraining, strongly suppressing verbatim recall of data while retaining downstream task performance. The Apertus models also expand multilingual coverage, training on 15T tokens from over 1800 languages, with ~40% of pretraining data allocated to non-English content. Released at 8B and 70B scales, Apertus approaches state-of-the-art results among fully open models on multilingual benchmarks, rivalling or surpassing open-weight counterparts. Beyond model weights, we release all scientific artifacts from our development cycle with a permissive license, including data preparation scripts, checkpoints, evaluation suites, and training code, enabling transparent audit and extension.

cs.CL

$O(a)$-improved QCD+QED Wilson Dirac operator on GPUs

Markov Chain Monte Carlo simulations of lattice Quantum Chromodynamics (QCD) are the only known tool to investigate non-perturbatively the theory of the strong interaction and are required to perform precision tests of the Standard Model of Particle Physics. As the Markov Chain is a serial process, the sole option for improving the sampling rate is accelerating each individual update step. Heterogeneous clusters of GPU-accelerated nodes offer large total memory bandwidth which can be used to speed-up our application, openQxD-1.1, which is dominated by inversions of the Dirac operator, a large sparse matrix. In this work we investigate offloading the inversion to GPU using the lattice-QCD library QUDA, and our early results demonstrate a significant potential speed-up in the time-to-solution for state-of-the-art problem sizes. Minimal extensions to the existing QUDA library are required for our specific physics programme while greatly enhancing the performance portability of our code and retaining the reliability and robustness of existing applications in openQxD-1.1. Our new interface will enable us to utilize pre-exascale infrastructure and reduce the systematic uncertainty in our physics predictions by incorporating the effects of quantum electromagnetism (QED) in our simulations.

hep-lat

NaCl salts in finite aqueous environments at the fine particle marine aerosol scale

We investigated isolated sodium/chloride aqueous droplets at the microscopic level, which comprise from about 5k to 1M water molecules and whose salt concentrations are 0.2$m$ (brackish water) and 0.6$m$ (sea water), by means of molecular dynamics simulations based on an \emph{ab initio}-based polarizable force field. The size of our largest droplets is at the submicron particle marine aerosol scale. From our simulations, we investigated ion spatial distributions, ion aggregates (size, composition, lifetime and distribution), droplet surface potentials and the densities of the water vapor surrounding the droplets. Regarding ions, they form a weak electrostatic double layer extending from the droplet boundary to 2~nm within the droplet interior. Free $\mathrm{Na^+}$ and ion aggregates are more repelled from the boundary than free $\mathrm{Cl^-}$. Most of the droplet properties depend on the droplet radius $R$ according to the standard formula $A=A_\infty(1 - 2 \delta/R) $, where $A_\infty$ is the bulk magnitude of the quantity $A$ and $\delta$ is a length at most at the~nm scale. Regarding the water vapor densities they obey a Kelvin relation corresponding to a surface tension whose Tolman length is negative and at the 1~nm scale. That length is about one order of magnitude larger than for pure water droplets, however it is weak enough to support the reliability of a standard Kelvin term (based on planar interface surface tensions and water densities) and of the related K{\"o}lher equation to model sub-micron salty aerosols.

physics.chem-ph

Models of gravitational lens candidates from Space Warps CFHTLS

We report modelling follow-up of recently-discovered gravitational-lens candidates in the Canada France Hawaii Telescope Legacy Survey. Lens modelling was done by a small group of specially-interested volunteers from the SpaceWarps citizen-science community who originally found the candidate lenses. Models are categorised according to seven diagnostics indicating (a) the image morphology and how clear or indistinct it is, (b) whether the mass map and synthetic lensed image appear to be plausible, and (c) how the lens-model mass compares with the stellar mass and the abundance-matched halo mass. The lensing masses range from ~10^11 Msun to >10^13 Msun. Preliminary estimates of the stellar masses show a smaller spread in stellar mass (except for two lenses): a factor of a few below or above ~10^11 Msun. Therefore, we expect the stellar-to-total mass fraction to decline sharply as lensing mass increases. The most massive system with a convincing model is J1434+522 (SW05). The two low-mass outliers are J0206-095 (SW19) and J2217+015 (SW42); if these two are indeed lenses, they probe an interesting regime of very low star-formation efficiency. Some improvements to the modelling software (SpaghettiLens), and discussion of strategies regarding scaling to future surveys with more and frequent discoveries, are included.

astro-ph.GA

Gravitational lens modelling in a citizen science context

We develop a method to enable collaborative modelling of gravitational lenses and lens candidates, that could be used by non-professional lens enthusiasts. It uses an existing free-form modelling program (glass), but enables the input to this code to be provided in a novel way, via a user-generated diagram that is essentially a sketch of an arrival-time surface. We report on an implementation of this method, SpaghettiLens, which has been tested in a modelling challenge using 29 simulated lenses drawn from a larger set created for the Space Warps citizen science strong lens search. We find that volunteers from this online community asserted the image parities and time ordering consistently in some lenses, but made errors in other lenses depending on the image morphology. While errors in image parity and time ordering lead to large errors in the mass distribution, the enclosed mass was found to be more robust: the model-derived Einstein radii found by the volunteers were consistent with those produced by one of the professional team, suggesting that given the appropriate tools, gravitational lens modelling is a data analysis activity that can be crowd-sourced to good effect. Ideas for improvement are discussed, these include (a) overcoming the tendency of the models to be shallower than the correct answer in test cases, leading to systematic overestimation of the Einstein radius by 10 per cent at present, and (b) detailed modelling of arcs.

astro-ph.IM

A Sampling Strategy for High-Dimensional Spaces Applied to Free-Form Gravitational Lensing

We present a novel proposal strategy for the Metropolis-Hastings algorithm designed to efficiently sample general convex polytopes in 100 or more dimensions. This improves upon previous sampling strategies used for free-form reconstruction of gravitational lenses, but is general enough to be applied to other fields. We have written a parallel implementation within the lens modeling framework GLASS. Testing shows that we are able to produce uniform uncorrelated random samples which are necessary for exploring the degeneracies inherent in lens reconstruction.

astro-ph.IM

Weak Microlensing

A nearby star having a near-transit of a galaxy will cause a time-dependent weak lensing of the galaxy. Because the effect is small, we refer to this as weak microlensing. This could provide a useful method to weigh low-mass stars and brown dwarfs. We examine the feasibility of measuring masses in this way and we find that a star causes measurable weak microlensing in a galaxy even at 10 Einstein radii away. Of order one magnitude I < 25 galaxy comes close enough to one or other of the ~100 nearest stars per year.

astro-ph.SR

A New Estimate of the Hubble Time with Improved Modeling of Gravitational Lenses

This paper examines free-form modeling of gravitational lenses using Bayesian ensembles of pixelated mass maps. The priors and algorithms from previous work are clarified and significant technical improvements are made. Lens reconstruction and Hubble Time recovery are tested using mock data from simple analytic models and recent galaxy-formation simulations. Finally, using published data, the Hubble Time is inferred through the simultaneous reconstruction of eleven time-delay lenses. The result is H_0^{-1}=13.7^{+1.8}_{-1.0} Gyr.

astro-ph