Fusing Complementary Multi-view Features for Screen-Based Eye Tracking
Current multi-view gaze estimation remains limited by existing datasets, insufficient exploitation of complementary cross-view information, and evaluation focused primarily on average gaze error. We address these limitations through a more systematic study of multi-view gaze estimation. First, we introduce PrismGaze, a new dataset with over three million images, capturing continuous headpose variation for the same gaze targets. Second, we propose PrismFusion, a multi-view feature fusion framework based on region partitioning, which masks complementary image regions across views during training to encourage effective cross-view information integration. Third, we develop a broader evaluation framework that examines the effects of camera number and placement, target location,and viewing depth. Our experiments show that the primary benefit of multi-view gaze estimation comes from compensating for poorly observed views with cameras providing more favorable viewpoints. PrismFusion remains robust to changes in viewing depth. Together, our dataset, method, and evaluation provide a more comprehensive foundation for studying multi-view gaze estimation in realistic settings.