arXiv · 2508.14685
SSA: Improving Performance With a Better Scoring Function
Abstract
While transformer models exhibit strong in-context learning (ICL) abilities, they often fail to generalize under simple distribution shifts. We analyze these failures and identify Softmax, the scoring function in the attention mechanism, as a contributing factor. We propose \textbf{Scaled Signed Averaging (SSA)}, a novel attention scoring function that mitigates these failures. SSA significantly improves performance on our ICL tasks and outperforms transformer models with Softmax on several NLP benchmarks and linguistic probing tasks, in both decoder-only and encoder-only architectures.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Omar Naim, Swarnadeep Bhar, Jérôme Bolte, Nicholas Asher. 2025-08-20. SSA: Improving Performance With a Better Scoring Function. https://arxiv.org/abs/2508.14685
Cite the original work for its findings. Save a collection to share your selection of sources.