Searcharxiv⌕ Search

arXiv subjects

Sophia N. Wilson

Publications and source records attributed to Sophia N. Wilson.

5 recordsLinked to original sources

Performance-Carbon Trade-Offs across Architectural Biases in Shear Flow Forecasting

Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics), leading to large-scale models that require massive computational resources and result in considerable carbon footprint across the model lifecycle. In this work, we explore how architectural biases, specifically a model's receptive field and periodicity assumption, are associated with the trade-offs between predictive performance and carbon footprint for spatio-temporal forecasting of incompressible shear flow. We study seven models differing in these properties and evaluate pointwise accuracy, physics-fidelity, and carbon cost across training and inference. We find that no single model dominates across both predictive performance and carbon cost, that a lower training cost does not straightforwardly extend to inference, and that runtime alone is an unreliable proxy for carbon cost, particularly during training. Together, these results underscore the importance of explicitly evaluating carbon costs across the full model lifecycle. We argue that model efficiency, alongside efficacy, should be a core consideration in machine learning model development and deployment.

cs.LG↗

How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI

Large-scale data has fuelled the success of frontier artificial intelligence (AI) models over the past decade. This expansion has relied on sustained efforts by large technology corporations to aggregate and curate internet-scale datasets. In this work, we examine the environmental, social, and economic costs of large-scale data in AI through a sustainability lens. We argue that the field is shifting from building models from data to actively creating data for building models. We characterise this transition as hyper-datafication, which marks a critical juncture for the future of frontier AI and its societal impacts. To quantify and contextualise data-related costs, we analyse approximately 550,000 datasets from the Hugging Face Hub, focusing on dataset growth, storage-related energy consumption and carbon footprint, and societal representation using language data. We complement this analysis with qualitative responses from data workers in Kenya to examine the labour involved, including direct employment by big tech corporations and exposure to graphic content. We further draw on external data sources to substantiate our findings by illustrating the global disparity in data centre infrastructure. Our analyses reveal that hyper-datafication drives substantial and growing environmental costs while systematically redistributing labour risks and representational harms toward the Global South. Thus, we propose Data PROOFS recommendations spanning provenance, resource awareness, ownership, openness, frugality, and standards to mitigate these costs. Our work aims to make visible the often-overlooked costs of data that underpin frontier AI and to stimulate broader debate within the research community and beyond.

cs.CY↗

Position: Stop Preaching and Start Practising Data Frugality for Responsible Development of AI

This position paper argues that the machine learning community must move from preaching to practising data frugality for responsible artificial intelligence (AI) development. For too long, progress has been equated with ever-larger datasets, driving remarkable advances but now yielding increasingly diminishing performance gains alongside rising energy use and carbon emissions. While awareness of data frugal approaches has grown, their adoption has remained rhetorical, and data scaling continues to dominate development practice. We argue that this gap between preach and practice must be closed, as continued data scaling entails substantial and under-accounted environmental impacts. To ground our position, we provide indicative estimates of the energy use and carbon emissions associated with the downstream use of ImageNet-1K. We then present empirical evidence that data frugality is both practical and beneficial, demonstrating that subset selection methods can substantially reduce training energy consumption with little loss in accuracy, while also mitigating dataset bias. Finally, we outline actionable recommendations for moving data frugality from rhetorical preaching to concrete practice for responsible development of AI.

cs.LG↗

Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis

Training large-scale deep neural networks (DNNs) is resource-intensive, making model compression a practical necessity. The widely accepted ''learning as compression'' hypothesis posits that training induces structure in network weights, which enables compression. Measuring this structure through Kolmogorov-Chaitin-Solomonoff (KCS) complexity is appealing, but existing estimators based on the Coding Theorem Method (CTM) and the Block Decomposition Method (BDM) are limited to small binary objects and do not scale to modern DNNs. We introduce the Quantized Block Decomposition method (QuBD), which extends algorithmic complexity estimation to any $k$-ary object. QuBD first quantizes the network weights to a finite alphabet, then estimates the KCS complexity by aggregating per bit-plane CTM estimates. We show theoretically that QuBD yields a strictly tighter estimation gap with respect to true KCS complexity than binarization-based methods. Using QuBD, we study how the algorithmic complexity of neural network weights evolves during training, showing that it decreases as models learn, scales with data budget, increases during overfitting, follows the delayed generalization observed during grokking, and correlates with generalization performance. We further show that algorithmic information resides predominantly in the most significant bit-planes, which can serve as a practical diagnostic for determining appropriate post-training quantization levels. This work offers novel insights into learning mechanisms in DNNs by providing the first scalable, tractable estimates of KCS complexity for large, non-binary objects such as DNN weights.

cs.LG↗

A high-redshift calibration of the [OI]-to-HI conversion factor in star-forming galaxies

The assembly and build-up of neutral atomic hydrogen (HI) in galaxies is one of the most fundamental processes in galaxy formation and evolution. Studying this process directly in the early universe is hindered by the weakness of the hyperfine 21-cm HI line transition, impeding direct detections and measurements of the HI gas masses ($M_{\rm HI}$). Here we present a new method to infer $M_{\rm HI}$ of high-redshift galaxies using neutral, atomic oxygen as a proxy. Specifically, we derive metallicity-dependent conversion factors relating the far-infrared [OI]-$63μ$m and [OI]-$145μ$m emission line luminosities and $M_{\rm HI}$ in star-forming galaxies at $z\approx 2-6$ using gamma-ray bursts (GRBs) as probes. We substantiate these results by observations of galaxies at $z\approx 0$ with direct measurements of $M_{\rm HI}$ and [OI]-$63μ$m and [OI]-$145μ$m in addition to hydrodynamical simulations at similar epochs. We find that the [OI]$_{\rm 63μm}$-to-HI and [OI]$_{\rm 145μm}$-to-HI conversion factors universally appears to be anti-correlated with the gas-phase metallicity. The high-redshift GRB measurements further predict a mean ratio of $L_{\rm [OI]-63μm} / L_{\rm [OI]-145μm}=1.55\pm 0.12$ and reveal generally less excited [CII]. The $z \approx 0$ galaxy sample also shows systematically higher $β_{\rm [OI]-63μm}$ and $β_{\rm [OI]-145μm}$ conversion factors than the GRB sample, indicating either suppressed [OI] emission in local galaxies or more extended, diffuse HI gas reservoirs traced by the HI 21-cm. Finally, we apply these empirical calibrations to the few high-redshift detections of [OI]-$63μ$m and [OI]-$145μ$m line transitions from the literature and further discuss the applicability of these conversion factors to probe the HI gas content in the dense, star-forming ISM of galaxies at $z\gtrsim 6$, well into the epoch of reionization.

astro-ph.GA↗