SearcharxivSearch

arXiv subjects

James Ball

Publications and source records attributed to James Ball.

9 recordsLinked to original sources

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings. However, how these models scale and how best to spend a pretraining budget remain poorly understood. We present the largest controlled scaling study for EO to date: 395 training runs within a fixed pixel-wise Barlow Twins family, each evaluated on 15 diverse downstream tasks. We find that pretraining loss barely predicts downstream performance (|Pearson r| < 0.2), so selecting models by loss wastes a large share of the compute. We also find that, as the training budget grows, the encoder and the data should grow together while the projector stays fixed, which gives a simple rule for allocating compute. Using this rule, we train a family of pixel-wise teachers (0.5B, 1B, and 2B) and distil the largest into compact students for embeddings-as-data deployment. In aggregate, our 44-million-parameter distilled student outperforms every open and proprietary embedding product we test, several of them an order of magnitude larger. These students produce Matryoshka representations that are inexpensive to serve: a 16-dimensional prefix keeps 92% of the full 128-dimensional performance at 1/8 of the storage. Together, these results give a concrete, empirically grounded recipe for scaling pixel-wise EO foundation models: train large encoders, select by downstream performance, and distil into flexible student models. We plan to release global 10 m annual embeddings covering 2017-2025 as version 2 of the TESSERA foundation-model embeddings product. All code is available at: https://github.com/ucam-eo/tessera

cs.CV

TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis

Satellite Earth-observation (EO) time series in the optical and microwave ranges of the electromagnetic spectrum are often irregular due to orbital patterns and cloud obstruction. Compositing addresses these issues but loses information with respect to vegetation phenology, which is critical for many downstream tasks. Instead, we present TESSERA, a pixel-wise foundation model for multi-modal (Sentinel-1/2) EO time series that learns robust, label-efficient embeddings. During model training, TESSERA uses Barlow Twins and sparse random temporal sampling to enforce invariance to the selection of valid observations. We employ two key regularizers: global shuffling to decorrelate spatial neighborhoods and mix-based regulation to improve invariance under extreme sparsity. We find that for diverse classification, segmentation, and regression tasks, TESSERA embeddings deliver state-of-the-art accuracy with high label efficiency, often requiring only a small task head and minimal computation. To democratize access, adhere to FAIR - principles, and simplify use, we release global, annual, 10m, pixel-wise int8 embeddings together with open weights/code and lightweight adaptation heads, thus providing practical tooling for large-scale retrieval and inference at planetary scale. All code and data are available at: https://github.com/ucam-eo/tessera.

cs.LG

URCDM: Ultra-Resolution Image Synthesis in Histopathology

Diagnosing medical conditions from histopathology data requires a thorough analysis across the various resolutions of Whole Slide Images (WSI). However, existing generative methods fail to consistently represent the hierarchical structure of WSIs due to a focus on high-fidelity patches. To tackle this, we propose Ultra-Resolution Cascaded Diffusion Models (URCDMs) which are capable of synthesising entire histopathology images at high resolutions whilst authentically capturing the details of both the underlying anatomy and pathology at all magnification levels. We evaluate our method on three separate datasets, consisting of brain, breast and kidney tissue, and surpass existing state-of-the-art multi-resolution models. Furthermore, an expert evaluation study was conducted, demonstrating that URCDMs consistently generate outputs across various resolutions that trained evaluators cannot distinguish from real images. All code and additional examples can be found on GitHub.

eess.IV

Ultra-Resolution Cascaded Diffusion Model for Gigapixel Image Synthesis in Histopathology

Diagnoses from histopathology images rely on information from both high and low resolutions of Whole Slide Images. Ultra-Resolution Cascaded Diffusion Models (URCDMs) allow for the synthesis of high-resolution images that are realistic at all magnification levels, focusing not only on fidelity but also on long-distance spatial coherency. Our model beats existing methods, improving the pFID-50k [2] score by 110.63 to 39.52 pFID-50k. Additionally, a human expert evaluation study was performed, reaching a weighted Mean Absolute Error (MAE) of 0.11 for the Lower Resolution Diffusion Models and a weighted MAE of 0.22 for the URCDM.

eess.IV

13 New Light Curves and Updated Mid-Transit Time and Period for Hot Jupiter WASP-104 b with EXOTIC

Using the EXOplanet Transit Interpretation Code (EXOTIC), we reduced 52 sets of images of WASP-104 b, a Hot Jupiter-class exoplanet orbiting WASP-104, in order to obtain an updated mid-transit time (ephemeris) and orbital period for the planet. We performed this reduction on images taken with a 6-inch telescope of the Center for Astrophysics | Harvard & Smithsonian MicroObservatory. Of the reduced light curves, 13 were of sufficient accuracy to be used in updating the ephemerides for WASP-104 b, meeting or exceeding the three-sigma standard for determining a significant detection. Our final mid-transit value was 2457805.170208 +/- 0.000036 BJD_TBD and the final period value was 1.75540644 +/- 0.00000016 days. The true significance of our results is in their derivation from image sets gathered over time by a small, ground-based telescope as part of the Exoplanet Watch citizen science initiative, and their competitive results to an ephemeris generated from data gathered by the TESS telescope. We use these results to further show how such techniques can be employed by amateur astronomers and citizen scientists to maximize the efficacy of larger telescopes by reducing the use of expensive observation time. The work done in the paper was accomplished as part of the first fully online Course-Based Undergraduate Research Experience (CURE) for astronomy majors in the only online Bachelor of Science program in Astronomical and Planetary Sciences.

astro-ph.EP

Realistic Data Enrichment for Robust Image Segmentation in Histopathology

Poor performance of quantitative analysis in histopathological Whole Slide Images (WSI) has been a significant obstacle in clinical practice. Annotating large-scale WSIs manually is a demanding and time-consuming task, unlikely to yield the expected results when used for fully supervised learning systems. Rarely observed disease patterns and large differences in object scales are difficult to model through conventional patient intake. Prior methods either fall back to direct disease classification, which only requires learning a few factors per image, or report on average image segmentation performance, which is highly biased towards majority observations. Geometric image augmentation is commonly used to improve robustness for average case predictions and to enrich limited datasets. So far no method provided sampling of a realistic posterior distribution to improve stability, e.g. for the segmentation of imbalanced objects within images. Therefore, we propose a new approach, based on diffusion models, which can enrich an imbalanced dataset with plausible examples from underrepresented groups by conditioning on segmentation maps. Our method can simply expand limited clinical datasets making them suitable to train machine learning pipelines, and provides an interpretable and human-controllable way of generating histopathology images that are indistinguishable from real ones to human experts. We validate our findings on two datasets, one from the public domain and one from a Kidney Transplant study.

cs.CV

Tree species classification from hyperspectral data using graph-regularized neural networks

We propose a novel graph-regularized neural network (GRNN) algorithm for tree species classification. The proposed algorithm encompasses superpixel-based segmentation for graph construction, a pixel-wise neural network classifier, and the label propagation technique to generate an accurate and realistic (emulating tree crowns) classification map on a sparsely annotated data set. GRNN outperforms several state-of-the-art techniques not only for the standard Indian Pines HSI but also achieves a high classification accuracy (approx. 92%) on a new HSI data set collected over the heterogeneous forests of French Guiana (FG) when less than 1% of the pixels are labeled. We further show that GRNN is competitive with the state-of-the-art semi-supervised methods and exhibits a small deviation in accuracy for different numbers of training samples and over repeated trials with randomly sampled labeled pixels for training.

cs.CV

Two-Orders-of-Magnitude Improvement in the Total Spin Angular Momentum of 131Xe Nuclei Using Spin Exchange Optical Pumping

We report on hyperpolarization of quadrupolar (I=3/2) 131Xe via spin-exchange optical pumping. Observations of the 131Xe polarization dynamics show that the effective alkali-metal/131Xe spin-exchange cross-sections are large enough to compete with 131Xe spin relaxation. 131Xe polarization up to 7.6 p/m 1.5 percent was achieved in ca. 8.5EE20 spins--a ca. 100-fold improvement in the total spin angular momentum--enabling applications including measurement of spin-dependent neutron-131Xe s-wave scattering and sensitive searches for time-reversal violation in neutron-131Xe interactions beyond the Standard Model.

physics.atom-ph

Critical Statistical Charge for Anyonic Superconductivity

We examine a criterion for the anyonic superconductivity at zero temperature in Abelian matter-coupled Chern-Simons gauge field theories in three dimensions. By solving the Dyson-Schwinger equations, we obtain a critical value of the statistical charge for the superconducting phase in a massless fermion-Chern-Simons model.

hep-th