Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate into better downstream performance. We investigate this disconnect using XLM-R models explicitly aligned across four language pairs with token-level, sentence-level, and masked-language-modeling objectives. We evaluate their zero-shot transfer on a token-level task (part-of-speech tagging) and a sentence-level task (sentence classification), and analyze both representational changes and the gradients induced by the alignment and downstream objectives. We find that embedding-based alignment metrics do not reliably indicate whether alignment will improve or degrade downstream performance. Moreover, alignment and downstream-task gradients are often nearly orthogonal, particularly when the alignment objective and downstream task operate at different representational levels. These findings suggest that representation alignment alone is insufficient for assessing cross-lingual transfer, and that the compatibility between alignment and downstream objectives should be considered when designing/evaluating alignment methods.