SiamJEPA: On the Role of Siamese Student Encoders in JEPA
Joint Embedding Predictive Architectures (JEPAs) have emerged as a promising framework for self-supervised representation learning by predicting latent embeddings of masked regions rather than reconstructing pixels. Existing JEPA methods typically employ a single student encoder, leaving the role of Siamese student encoders largely unexplored. In this paper, we propose Siamese JEPA (SiamJEPA), a JEPA framework with masked Siamese student encoders and an exponential moving average (EMA) teacher, which can also be viewed as a JEPA formulation of the brain-inspired representation learning model PhiNet. We further introduce Random Shuffle Teacher (RST), which removes spatial correspondence in teacher targets to encourage semantic patch representations, and develop an RST-based semantic-to-spatial curriculum. Experiments on ImageNet show that Siamese student encoders effectively regularize the JEPA objective, improving representation separability and accelerating early-stage learning. Moreover, under RST, stronger Siamese regularization substantially increases class-discriminative information in individual patch tokens, suggesting that the semantic bias arises from the interaction between RST and the Siamese objective rather than from shuffling alone. SiamJEPA consistently outperforms comparable single-encoder JEPA variants under limited training budgets. With RST-based curriculum learning and a ViT-Base backbone, SiamJEPA achieves 74.2\% linear-probing accuracy after 450 epochs using a substantially simpler masking strategy, compared with I-JEPA (72.9\%) and DSeq-JEPA (73.5\%) after 600 epochs. These results demonstrate that Siamese student encoders provide an effective inductive bias for predictive representation learning and can be further enhanced through semantic-to-spatial curriculum learning. The source code is publicly available at https://github.com/oist/SiamJEPA.