arXiv · 2609.31764
Oracle Complementarity Is Not Realizable Complementarity in Frozen-Encoder Audio-Visual Emotion Recognition
Abstract
Complementarity analyses of multimodal systems commonly report an oracle ceiling (the fraction of examples on which at least one unimodal branch is correct) and interpret the gap between it and realized fusion accuracy as recoverable headroom. We show that this ceiling is not a fusion target. Using frozen self-supervised audio and visual encoders with a trained head, we measure the ceiling and realized gain on CREMA-D under two audio encoders (one whose fine-tuning lineage includes CREMA-D, one clean) and on EAV under the contaminated encoder. Replacing the fine-tuned encoder with a self-supervised one triples the headroom, from +0.055to +0.186. A second clean encoder and a controlled degradation of the contaminated one fall on the same curve, so provenance moves the headroom by changing audio-branch accuracy. Concatenation realizes at most about half of the headroom (converted fraction 0.55), and on EAV the converted fraction is indistinguishable from zero despite larger headroom. A learned router is beaten by plain concatenation everywhere. Oracle headroom measures branch disagreement, not a fusion budget.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Benjamin Hurt. 2026-09-24. Oracle Complementarity Is Not Realizable Complementarity in Frozen-Encoder Audio-Visual Emotion Recognition. https://arxiv.org/abs/2609.31764
Cite the original work for its findings. Save a collection to share your selection of sources.