SearcharxivSearch

arXiv subjects

Stephen J. Thomas

Publications and source records attributed to Stephen J. Thomas.

5 recordsLinked to original sources

Gated Subspace Inference for Transformer Acceleration

A method is presented for accelerating inference in transformer language models by exploiting the low effective rank of the token activation manifold at each layer. The method decomposes each activation vector into a subspace component and a residual, computes the linear-layer output on the subspace component via a cached low-rank weight image at reduced memory bandwidth, and applies a per-token gate that determines whether the residual correction is computed or skipped. The gate ensures that the output distribution is preserved to within a controllable tolerance. Validation on three model families (GPT-2 124M, GPT-J 6B, OPT 6.7B) on AMD MI300X demonstrates effective speedups of 3.0x to 10.5x on linear-layer weight reads with perplexity ratios below 1.00 and top-1 token agreement above 98%. The method requires no retraining, no architectural modification, and no approximation of the attention mechanism. At the operating point (k = 256, ε = 0.05) on GPT-J 14 6B, the accelerated model produces character-for-character identical output to the baseline.

cs.LG

Cascade Token Selection for Transformer Attention Acceleration

A method is presented for reducing the cost of representative token selection in transformer attention layers by exploiting the coherence of the representative set across depth. Activation Decorrelation Attention (ADA) selects $r \ll T$ representative tokens at each layer via a Gram threshold and computes attention on the compressed $r \times r$ problem, but the selection requires a $T \times T$ Gram matrix at every layer. The cascade mechanism introduced here inherits the representative set from layer $l$ to layer $l+1$, validates it via a $(T - r) \times r$ cross-Gram computation, and updates it with a small number of additions and removals. The cost of the selection step drops from $O(T^2 d)$ to $O(T r d)$ per layer. Validation on three model families (GPT-2 124M, GPT-J 6B, OPT 6.7B) on AMD MI300X demonstrates Gram operation savings of $22\%$ to $63\%$ with mean Jaccard overlap of $0.83$ to $0.94$ between consecutive layers. The cascade reveals that the set of informative tokens is a structural property of the input that propagates coherently through the depth of the network: the same tokens carry the non-redundant information at layer $l$ and at layer $l+1$.

cs.LG

Linear solvers for power grid optimization problems: a review of GPU-accelerated linear solvers

The linear equations that arise in interior methods for constrained optimization are sparse symmetric indefinite and become extremely ill-conditioned as the interior method converges. These linear systems present a challenge for existing solver frameworks based on sparse LU or LDL^T decompositions. We benchmark five well known direct linear solver packages using matrices extracted from power grid optimization problems. The achieved solution accuracy varies greatly among the packages. None of the tested packages delivers significant GPU acceleration for our test cases.

math.NA

A Comparison of Two Shallow Water Models with Non-Conforming Adaptive Grids: classical tests

In an effort to study the applicability of adaptive mesh refinement (AMR) techniques to atmospheric models an interpolation-based spectral element shallow water model on a cubed-sphere grid is compared to a block-structured finite volume method in latitude-longitude geometry. Both models utilize a non-conforming adaptation approach which doubles the resolution at fine-coarse mesh interfaces. The underlying AMR libraries are quad-tree based and ensure that neighboring regions can only differ by one refinement level. The models are compared via selected test cases from a standard test suite for the shallow water equations. They include the advection of a cosine bell, a steady-state geostrophic flow, a flow over an idealized mountain and a Rossby-Haurwitz wave. Both static and dynamics adaptations are evaluated which reveal the strengths and weaknesses of the AMR techniques. Overall, the AMR simulations show that both models successfully place static and dynamic adaptations in local regions without requiring a fine grid in the global domain. The adaptive grids reliably track features of interests without visible distortions or noise at mesh interfaces. Simple threshold adaptation criteria for the geopotential height and the relative vorticity are assessed.

physics.comp-ph

Multiplicative, Additive and Restricted Additive Optimized Schwarz Preconditioning

We demonstrate that a small modification of the multiplicative, additive and restricted additive Schwarz preconditioner at the algebraic level, motivated by optimized Schwarz methods defined at the continuous level, leads to a significant reduction in the iteration count of the iterative solver. Numerical experiments using finite difference and spectral element discretizations of the positive definite Helmholtz problem and an idealized climate simulation illustrate the effectiveness of the new approach.

math.NA