arXiv · 2609.28117
Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation
Abstract
In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps. This framework enables a large-scale causal analysis of attention heads, making it suitable for LLMs. We evaluate our method on the task of disambiguation in Context-aware Machine Translation, where we analyze 50 phenomena across 4 models and 4 language directions. We empirically show the alignment of our method with the effects of increasing the attention scores of token-to-token relations on three models and two language directions, ensuring the robustness of our method. Our analysis reveals the presence of the "general-purpose" attention heads that improve the model's performance when attending to different relations. We find that the average attention a head assigns to a relation does not necessarily relate to the model's performance, which suggests that the models developed redundancies during training in terms of the head functions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis. 2026-09-23. Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation. https://arxiv.org/abs/2609.28117
Cite the original work for its findings. Save a collection to share your selection of sources.