arXiv · 2608.27760
Informational Antilocality and the Locality Bias in LLMs
Abstract
We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Andrew McInnerney, Shane Storks, Steven Abney, Richard L. Lewis. 2026-08-27. Informational Antilocality and the Locality Bias in LLMs. https://arxiv.org/abs/2608.27760
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.