SearcharxivSearch

arXiv subjects

Joe Dwyer

Publications and source records attributed to Joe Dwyer.

5 recordsLinked to original sources

A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget

This study examines training dynamics in a small Llama-style language model trained under a fixed, compute-constrained token budget. Rather than evaluating efficiency solely through endpoint performance, the study uses a quantitative experimental repeated measures design to analyze how validation loss, validation perplexity, rolling volatility, backslide behavior, spike behavior, and between-seed variability change across token-based training intervals. Six independent training runs were conducted on a 4.26-million-parameter model using the TinyStories corpus, CPU-based full-precision training, and a target budget of approximately 20 million cumulative training tokens. Metrics were collected across 21 intervals, producing 126 seed-by-interval observations. Repeated measures ANOVA showed statistically significant interval effects for validation loss, validation perplexity, and rolling volatility. Descriptive trajectories revealed rapid early improvement followed by non-monotonic degradation during later training intervals. Mean validation loss decreased from 8.3552 at initialization to 2.7996 near 4 million tokens, but increased to 3.9010 by the final checkpoint. Validation perplexity followed the same pattern, falling sharply early in training before rising later. Derived telemetry further showed recurrent validation-loss backslides and no interval-summary evidence of a stable phase under the predefined criteria. These findings suggest that compute-aware language model evaluation should examine training trajectories rather than endpoint metrics alone. In constrained compute settings, additional token exposure may increase computational cost without producing proportional generalization gains, and interval-level telemetry can reveal instability, regression, and diminishing returns that final metrics may obscure.

cs.AI

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

Research in machine learning has questioned whether increases in training token counts reliably produce proportional performance gains in large language models. Building on prior work introducing an energy-aware parameter efficiency metric, this study empirically examines the effects of increasing training token counts under fixed hardware and training conditions. The significance of this work lies in the explicit integration of power consumption and execution duration, as reflected by the power sampling frequency, into token-scale analysis. This addresses a gap in prior studies emphasizing performance outcomes while underrepresenting computational and energy costs. Using a repeated-measures experimental design on a constant GPU instance with an identical model architecture, optimizer settings, and epoch counts, a 1.1-billion-parameter TinyLlama model was trained at three token counts (500K, 1M, and 2M). While conventional performance metrics exhibited inconsistent or diminishing returns across token scales, the inclusion of power consumption and execution duration revealed a strictly monotonic decline in training efficiency as token count increased. Repeated-measures ANOVA demonstrated a strong effect of token count on parameter efficiency, with all pairwise comparisons remaining significant following Bonferroni correction. These findings indicate that increases in training token counts may be energetically inefficient even when marginal performance improvements are observed, underscoring the importance of efficiency-aware evaluation in large language model training.

cs.LG

Measuring the locations and properties of VHF sources emitted from an aircraft flying through high clouds

We show that it is possible to locate the few places on the body of an airplane, while it is flying through high clouds, from which broad-band, pulsed, radiation is emitted at Very High Frequency (VHF) radio frequencies. This serendipitous discovery was made whilst imaging a lightning flash using the Low-Frequency Array (LOFAR). This observation provides insights into the way the airplane sheds the electrical charge it acquires when flying through clouds. Furthermore, this observation allowed us to test and improve the precision and accuracy for our lightning observation techniques. Our new results indicate that with the improved procedure the location precision for strong pulses is better than 50~cm, with the orientation of linear polarization being accurate to within 25$^\circ$. For the present case of a Boeing 777-300ER, VHF emissions were observed exclusively associated with the two engines, as well as a specific spot on the tail. Despite the aircraft flying through clouds at an altitude of 8~km, we did not detect any emissions from electrostatic wicks.

physics.plasm-ph

Interferometric imaging of Intensely Radiating Negative Leaders

The common phenomenon of lightning still harbors many secrets and only recently a new propagation mode was observed for negative leaders. While propagating in this `Intensely Radiating Negative Leader' (IRNL) mode a negative leader emits 100 times more very-high frequency (VHF) and broadband radiation than a more normal negative leader. We have reported that this mode occurs soon after initiation of all lightning flashes we have mapped as well as sometimes long thereafter. Because of the profuse emission of VHF the leader structure is very difficult to image. In this work we report on measurements made with the LOFAR radio telescope, an instrument primarily built for radio-astronomy observations. For this reason, as part of the present work, we have refined our time resolved interferometric 3-Dimensional (TRI-D) imaging to take into account the antenna function. The images from the TRI-D imager show that during an IRNL there is an ionization front with a diameter in excess of 500~m where strong corona bursts occur. This is very different from what is seen for a normal negative leader where the corona bursts happen at the tip, an area of typically 10~m in diameter. The observed massive ionization wave supports the idea that this mode is indicative of a dense charge pocket.

physics.ao-ph

Time resolved 3Dinterferometric imaging of a section of a negative leader with LOFAR

We have developed a three dimensional (3D) interferometric beamforming technique for imaging lightning flashes using Very-High Frequency (VHF) radio data recorded from several hundreds antennas with baselines up to 100~km as offered by the Low Frequency Array (LOFAR). The long baselines allow us to distinguish fine structures on the scale of meters while the large number of antennas allow us to observe processes that radiate at the same intensity as the background when using a time resolution that is close to the impulse-response time of the system, 100~ns. The new beamforming imaging technique is complementary to our existing impulsive imaging technique. We apply this new tool to the imaging of a four stepped negative leaders in two flashes. For one flash, we observe the dynamics of coronal flashes that are emitted in the stepping process. Additionally, we show that the intensity emitted in VHF during the stepping process follows a power-law over 4 orders of magnitude in intensity for four leaders in two different lightning storms.

physics.ao-ph