arXiv · 2610.05463
Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs
Abstract
Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates ($ρ= 0.82$), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Renchi Zhang, Joost C. F. de Winter, Dimitra Dodou, Harleigh C. Seyffert, Yke Bauke Eisma. 2026-10-04. Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs. https://doi.org/10.1007/s42113-026-00333-4
Cite the original work for its findings. Save a collection to share your selection of sources.