Vision-Language Model Ensembles Achieve Human-Expert Accuracy for Galaxy Merger Classification
Models (VLMs) combined using a Bayesian statistical framework can classify galaxy merger morphologies with accuracy comparable to trained human experts. We deploy 15 VLM classifier configurations, spanning four model architectures (Gemma-4 E2B, Gemma-4 E4B, Qwen2.5-VL, and Qwen3-VL) tested with up to four prompt engineering strategies each. We evaluate their performance against a truth-known sample of 41 VELA+SUNRISE mock galaxy images from Lambrides et al. (2021a). As a proof-of-concept, all validation is performed on these mock images; application to real observational samples will require additional calibration and observational validation. The VLM ensemble achieves 83.3% accuracy on confident classifications (merger probability pM >= 0.8 or pM <= 0.2) and 58.3% completeness, with 5 misclassified galaxies, compared to 85% for both human accuracy and completeness. The ensemble recovers the population merger fraction to within 0.66 sigma of the truth (fM = 0.52 +/- 0.09 vs. true value of 0.585). Bayesian weighting improves overall accuracy by 17.1 percentage points over simple majority voting, with sensitivity improving by 29.2 percentage points. The ensemble produces 5 misclassifications (2 false positives, 3 false negatives), comparable to the 6 misclassifications (5 false positives, 1 false negative) reported for human classifiers by L21. The error-profile differences are not statistically significant for this sample. VLMs also produce more moderate per-galaxy merger probability distributions (27% uncertain) than the more polarized human distributions (15% uncertain), though this difference is also consistent with statistical fluctuation. These results establish VLMs as scalable, reproducible alternatives to human classifiers within a Bayesian probabilistic merger-fraction framework, for large-survey applications.