SearcharxivSearch

arXiv subjects

Peter Carson

Publications and source records attributed to Peter Carson.

3 recordsLinked to original sources

Scaling Inference Prefill with High-Radix Photonic Interconnects

With the rise of inference as today's dominant AI workload, the industry is transitioning to high-bandwidth photonic interconnects to meet the large scale-up requirements of increasingly complex Mixture-of-Experts (MoE) models. This paper quantifies the benefits of 3D-integrated photonic interconnects for inference prefill by analyzing tradeoffs between high-concurrency throughput for Large Language Model (LLM) chat and the large context windows typically required for reasoning and agentic AI. We simulate three MoE models: short context (1K--8K tokens), medium context (128K tokens), and long context (1M tokens). We evaluate this workload across existing copper-based GPU systems and one with high bandwidth integrated photonics. We show 2.1--3.2x latency improvements in the stressed high-batch regimes and 2.8--5.8x improvements over baselines in communication-limited configurations. 3D photonics enable the 1152-GPU footprint required to lower time-to-first-token, yielding 2.2--4.5x speedups across production-grade platforms when electrical systems cross their inherent scale-up-pod limits.

cs.DC

Long-term outburst activity of comet 17P/Holmes and constraints on ejecta size distributions

A quantitative understanding of cometary outbursts requires robust constraints on the size distribution of ejected particles, which governs outburst dynamics and underpins estimates of released gas and dust. In the absence of direct measurements of particle sizes, assumptions about the size distribution play a central role in modelling dust-trail formation, their dynamical evolution and observability, and the potential production of meteor showers following encounters with Earth. We analyse brightness amplitude variations associated with outbursts of comet 17P/Holmes from 1892 to 2021, with particular emphasis on the exceptional 2007 mega-outburst. During this event the comet underwent a rapid and substantial brightening: at its peak, the expanding coma reached a diameter exceeding that of the Sun and briefly became the largest object in the Solar System visible to the naked eye. We constrain the size distribution and total mass of porous agglomerates composed of ice, organics, and dust ejected during the outburst. The inferred particle size distribution is consistent with a power law of index q, yielding effective particle sizes between 10^-6 m for q = 4 and 5 x 10^-3 m for q = 2. Accounting for effective particle size, sublimation flux, and bulk density, we find that the total number of ejected particles increases with both q and sublimation flux. These results place constraints on the physical properties of outburst ejecta and provide physically motivated initial conditions for long-term dust-trail evolution modelling. They further indicate that cometary outburst brightness is determined primarily by the number of particles and their size distribution, rather than by the total ejected mass alone, with direct implications for the origin and evolution of meteoroid streams and the interplanetary dust population.

astro-ph.EP

Accelerating Frontier MoE Training with 3D Integrated Optics

The unabated growth in AI workload demands is driving the need for concerted advances in compute, memory, and interconnect performance. As traditional semiconductor scaling slows, high-speed interconnects have emerged as the new scaling engine, enabling the creation of larger logical GPUs by linking many GPUs into a single, low-latency, high-bandwidth compute domain. While initial scale-up fabrics leveraged copper interconnects for their power and cost advantages, the maximum reach of passive electrical interconnects (approximately 1 meter) effectively limits the scale-up domain to within a single rack. The advent of 3D-stacked optics and logic offers a transformative, power-efficient scale-up solution for connecting hundreds of GPU packages (thousands of GPUs) across multiple data center racks. This work explores the design tradeoffs of scale-up technologies and demonstrates how frontier LLMs necessitate novel photonic solutions to achieve aggressive power and performance targets. We model the benefits of 3D CPO (Passage) enabled GPUs and switches within the scale-up domain when training Frontier Mixture of Experts (MoE) models exceeding one trillion parameters. Our results show that the substantial increases in bandwidth and radix enabled by 3D CPO allow for an 8X increase in scale-up capability. This affords new opportunities for multi-dimensional parallelism within the scale-up domain and results in a 2.7X reduction in time-to-train, unlocking unprecedented model scaling.

cs.AR