arXiv · 2610.06102
Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation
Abstract
We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without task-specific training. For periocular, zero-shot predictions are strongly biased towards males, primarily due to a misaligned decision boundary. Threshold alignment substantially reduces this bias, reaching 85.29% accuracy. Linear SVMs trained on CLIP features provide only marginal gains, with a best periocular accuracy of 86.17%, approximately 2.8% above previous Adience results in the literature. Nevertheless, the gap with full-face performance confirms the greater difficulty of periocular gender estimation
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fernando Alonso-Fernandez, Kevin Hernandez-Diaz, Jose Maria Buades, Josef Bigun. 2026-10-05. Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation. https://arxiv.org/abs/2610.06102
Cite the original work for its findings. Save a collection to share your selection of sources.