arXiv · 2609.05074
Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
Abstract
We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.
Explore related subjects
Keep this discovery
Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi. 2026-09-04. Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time. https://arxiv.org/abs/2609.05074
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.