SearcharxivSearch

arXiv subjects

Nestor Demeure

Publications and source records attributed to Nestor Demeure.

3 recordsLinked to original sources

GIGA-Lens 2.0: Strong-Lens Modeling on Multiple GPU Nodes

We present GIGA-Lens 2.0: a major upgrade to the GPU-accelerated Bayesian framework for modeling strong lensing systems that allows it to be run across multiple GPU nodes. We have succeeded in running GIGA-Lens 2.0 on 128 nodes or 512 A100 GPUs. We demonstrate the speed benefits of this new version, and apply them to modeling 100 simulated systems and a real system, DESI J238.5690+04.7276. We also present other changes to the framework that have yielded further improvement on performance.

astro-ph.CO

Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores

New large language models (LLMs) are being released every day. Some perform significantly better or worse than expected given their parameter count. Therefore, there is a need for a method to independently evaluate models. The current best way to evaluate a model is to measure its Elo score by comparing it to other models in a series of contests - an expensive operation since humans are ideally required to compare LLM outputs. We observe that when an LLM is asked to judge such contests, the consistency with which it selects a model as the best in a matchup produces a metric that is 91% correlated with its own human-produced Elo score. This provides a simple proxy for Elo scores that can be computed cheaply, without any human data or prior knowledge.

cs.AI

Ranger21: a synergistic deep learning optimizer

As optimizers are critical to the performances of neural networks, every year a large number of papers innovating on the subject are published. However, while most of these publications provide incremental improvements to existing algorithms, they tend to be presented as new optimizers rather than composable algorithms. Thus, many worthwhile improvements are rarely seen out of their initial publication. Taking advantage of this untapped potential, we introduce Ranger21, a new optimizer which combines AdamW with eight components, carefully selected after reviewing and testing ideas from the literature. We found that the resulting optimizer provides significantly improved validation accuracy and training speed, smoother training curves, and is even able to train a ResNet50 on ImageNet2012 without Batch Normalization layers. A problem on which AdamW stays systematically stuck in a bad initial state.

cs.LG