ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs
Reliable visual recognition of shellfish is important for ecological monitoring and aquaculture, but real-world imagery exhibits substantial variation in acquisition conditions, appearance, and posture. Existing marine vision resources cover diverse taxa and imaging settings, but fine-grained mollusc recognition remains underexplored in a group-disjoint benchmark that jointly evaluates class prediction, representation-space retrieval, and condition sensitivity. To address this gap, we introduce ShellfishNet, a provenance-aware, group-disjoint benchmark of 40,130 images from 100 marine mollusc classes, organized into 30,879 acquisition or observation groups. Rather than treating recognition as a single clean-accuracy leaderboard, ShellfishNet evaluates model configurations across classification, image-to-image retrieval, taxonomic error structure, and controlled condition sensitivity. We benchmark representative convolutional, transformer-based, hybrid and efficient, self-supervised, and vision-language models on classification and retrieval. Separately, we evaluate open image-captioning and vision-language models on the captioning subset and assess selected fixed checkpoints under controlled generic and marine-inspired synthetic transformations. Across the benchmark, image-level and class-aware classification largely agree, whereas retrieval provides complementary model rankings and taxonomy projection reveals that a non-negligible subset of exact-class errors remains correct at higher taxonomic ranks. Selected fixed checkpoints with similar clean accuracy can also exhibit distinct sensitivity profiles under controlled transformations. These results show that model comparison on shellfish imagery should account for both task and image condition, rather than rely on a single clean classification score. The dataset and evaluation protocols will be publicly released.