SearcharxivSearch

arXiv subjects

Shangkun Wang

Publications and source records attributed to Shangkun Wang.

8 recordsLinked to original sources

MaxKernel: Agentic Kernel Generation for TPUs

Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic systems for kernel generation. In this work, we present MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: (1) a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; (2) an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and (3) a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space. All three paradigms leverage a shared pool of specialized sub-agents to handle planning, implementation, self-debugging, testing, and hardware profiling. We evaluate MaxKernel on JaxBench, a comprehensive suite of 50 diverse kernel tasks for TPUs, alongside complex, real-world workloads from state-of-the-art open-source models. We demonstrate that MaxKernel consistently generates highly optimized implementations, matching expert hand-tuned baselines and delivering significant performance across the benchmark. Our agent is open-sourced and available https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel.

cs.AI

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We present JAXBench, a TPU-native benchmark suite for AI-generated kernel optimization on Google Cloud TPUs. JAXBench comprises 50 JAX workloads that are both relevant and provide headroom for optimization. We extract 17 production ML operators from architectures in the public MaxText library such as Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, and AlphaFold2, and translate 33 operators from KernelBench that are validated for correctness and set with new problem sizes that achieve high TPU v6e MXU utilization. Eight of the 17 production operators ship with hand-optimized Pallas kernels from the public Tokamax library and block-size tuned to establish an expert upper-bound baseline. We evaluate four feedback-driven methods on generating candidate Pallas kernels for JAXBench. Across the full suite with Gemini 3 Flash, we find that target-specific context matters more than model scale on a sparsely-documented DSL like Pallas. Conditioning on curated TPU documentation raises per-sample correctness from 5.8% to 37.3% and solves 48 of 50 benchmarks at a 1.28x geomean speedup. Search structure yields significant gains once correctness is achieved, with Autocomp's beam-search pipeline reaching a 1.36x geomean speedup over XLA. On the 8 hand-tuned kernels, Autocomp reaches 1.60x geomean over XLA, recovering most of the 2.08x Tokamax upper bound but trailing on the specialized paged and ragged attention operators. High-quality TPU kernel optimization remains a challenging task, and we release the JAXBench benchmark, evaluation harness, and baseline results to support open source contributions.

cs.AI

Active Learning via Heteroskedastic Rational Kriging

Active learning methods for emulating complex computer models that rely on stationary Gaussian processes tend to produce design points that uniformly fill the entire experimental region, which can be wasteful for functions which vary only in small regions. In this article, we propose a new Gaussian process model that captures the heteroskedasticity of the function. Active learning using this new model can place design points in the more interesting regions of the response surface, and thus obtain surrogate models with better accuracy. The proposed active learning method is compared with the state-of-the-art methods using simulations and two real datasets. It is found to have comparable or better performance relative to other non-stationary Gaussian process-based methods, but faster by orders of magnitude.

stat.ME

Asset Bundling for Wind Power Forecasting

The growing penetration of intermittent, renewable generation in US power grids, especially wind and solar generation, results in increased operational uncertainty. In that context, accurate forecasts are critical, especially for wind generation, which exhibits large variability and is historically harder to predict. To overcome this challenge, this work proposes a novel Bundle-Predict-Reconcile (BPR) framework that integrates asset bundling, machine learning, and forecast reconciliation techniques. The BPR framework first learns an intermediate hierarchy level (the bundles), then predicts wind power at the asset, bundle, and fleet level, and finally reconciles all forecasts to ensure consistency. This approach effectively introduces an auxiliary learning task (predicting the bundle-level time series) to help the main learning tasks. The paper also introduces new asset-bundling criteria that capture the spatio-temporal dynamics of wind power time series. Extensive numerical experiments are conducted on an industry-size dataset of 283 wind farms in the MISO footprint. The experiments consider short-term and day-ahead forecasts, and evaluates a large variety of forecasting models that include weather predictions as covariates. The results demonstrate the benefits of BPR, which consistently and significantly improves forecast accuracy over baselines, especially at the fleet level.

stat.ME

Sequential Designs for Filling Output Spaces

Space-filling designs are commonly used in computer experiments to fill the space of inputs so that the input-output relationship can be accurately estimated. However, in certain applications such as inverse design or feature-based modeling, the aim is to fill the response or feature space. In this article, we propose a new experimental design framework that aims to fill the space of the outputs (responses or features). The design is adaptive and model-free, and therefore is expected to be robust to different kinds of modeling choices and input-output relationships. Several examples are given to show the advantages of the proposed method over the traditional input space-filling designs.

stat.ME

Risk-Aware Control and Optimization for High-Renewable Power Grids

The transition of the electrical power grid from fossil fuels to renewable sources of energy raises fundamental challenges to the market-clearing algorithms that drive its operations. Indeed, the increased stochasticity in load and the volatility of renewable energy sources have led to significant increases in prediction errors, affecting the reliability and efficiency of existing deterministic optimization models. The RAMC project was initiated to investigate how to move from this deterministic setting into a risk-aware framework where uncertainty is quantified explicitly and incorporated in the market-clearing optimizations. Risk-aware market-clearing raises challenges on its own, primarily from a computational standpoint. This paper reviews how RAMC approaches risk-aware market clearing and presents some of its innovations in uncertainty quantification, optimization, and machine learning. Experimental results on real networks are presented.

math.OC

Quasi-Ballistic Thermal Conduction in 6H-SiC

The minimization of electronics makes heat dissipation of related devices an increasing challenge. When the size of materials is smaller than the phonon mean free paths, phonons transport without internal scatterings and laws of diffusive thermal conduction fail, resulting in significant reduction in the effective thermal conductivity. This work reports, for the first time, the temperature dependent thermal conductivity of doped epitaxial 6H-SiC and monocrystalline porous 6H-SiC below room temperature probed by time-domain thermoreflectance. Strong quasi-ballistic thermal transport was observed in these samples, especially at low temperatures. Doping and structural boundaries were applied to tune the quasi-ballistic thermal transport since dopants selectively scatter high-frequency phonons while boundaries scatter phonons with long mean free paths. Exceptionally strong phonon scattering by boron dopants are observed, compared to nitrogen dopants. Furthermore, orders of magnitude reduction in the measured thermal conductivity was observed at low temperatures for the porous 6H-SiC compared to the epitaxial 6H-SiC. Finally, first principles calculations and a simple Callaway model are built to understand the measured thermal conductivities. Our work sheds light on the fundamental understanding of thermal conduction in technologically-important wide bandgap semiconductors such as 6H-SiC and will impact applications such as thermal management of 6H-SiC-related electronics and devices.

cond-mat.mtrl-sci

Deformation characteristics of a single droplet driven by a piezoelectric nozzle of the drop-on-demand inkjet system

In the drop-on-demand (DOD) inkjet system, deformation process and the direct relations between the droplet motions and the liquid properties have been seldom investigated, although they are very critical for the printing accuracy. In this study, experiments and computational simulation regarding deformation of a single droplet driven by a piezoelectric nozzle have been conducted to address the deformation characteristics of droplets. It is found that the droplet deformation is influenced by the pressure wave propagation in the ink channel related to the driven parameters and reflected in the subsequent droplet motions. The deformation extent oscillates with a certain period of T and a decreasing amplitude as the droplet moves downwards. The deformation extent is found strongly dependent on the capillary number (Ca), first ascended and then descended as the number increases. The maximum value of the deformation extent is surprisingly found to be within range of 0.68-0.82 of the Ca number regardless of other factors. Furthermore, the Rayleigh's linear relation of the oscillation frequency of the droplet to the parameter,({\sigma}/(\r{ho}r^3))^(1/2), is updated with a smaller slope shown both by experiments and simulation.

physics.flu-dyn