arXiv · 2310.07325
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
Abstract
Prior work suggests that language models manage the limited bandwidth of the residual stream through a "memory management" mechanism, where certain attention heads and MLP layers clear residual stream directions set by earlier layers. Our study provides concrete evidence for this erasure phenomenon in a 4-layer transformer, identifying heads that consistently remove the output of earlier heads. We further demonstrate that direct logit attribution (DLA), a common technique for interpreting the output of intermediate transformer layers, can show misleading results by not accounting for erasure.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jett Janiak, Can Rager, James Dao, Yeu-Tong Lau. 2023-10-11. An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L. https://doi.org/10.18653/v1%2F2024.blackboxnlp-1.15
Cite the original work for its findings. Save a collection to share your selection of sources.