arXiv · 2310.07325
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
Abstract
Prior work suggests that language models manage the limited bandwidth of the residual stream through a "memory management" mechanism, where certain attention heads and MLP layers clear residual stream directions set by earlier layers. Our study provides concrete evidence for this erasure phenomenon in a 4-layer transformer, identifying heads that consistently remove the output of earlier heads. We further demonstrate that direct logit attribution (DLA), a common technique for interpreting the output of intermediate transformer layers, can show misleading results by not accounting for erasure.
Explore related subjects
Keep this discovery
Jett Janiak, Can Rager, James Dao, Yeu-Tong Lau. 2023-10-11. An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L. https://doi.org/10.18653/v1%2F2024.blackboxnlp-1.15
Cite the original work for its findings. Save a collection to share your selection of sources.