arXiv · 2609.24026
InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
Abstract
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between adjacent super-/sub-categories. However, such fixed connectors may not optimally capture the relationships within a semantic hierarchy. To address this limitation, we propose interconnected hierarchical semantic representations (InterHier), which utilize a prepended learnable context to globally guide the interpretation of prompts containing hierarchical relationships. InterHier operates in two main stages. First, it constructs a hierarchy-aware prompt by integrating super-/sub-categories and prepending a learnable context. Second, it optimizes this learnable context to align visual region embeddings and textual embeddings. InterHier consistently improves performance over methods that rely on fixed connectors and can be seamlessly integrated into existing open-vocabulary object detection models. Experiments on open-vocabulary object detection benchmarks demonstrate that InterHier achieves competitive performance against state-of-the-art methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yeong-Jin Kim, Ho-Joong Kim, Seong-Whan Lee. 2026-09-21. InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection. https://doi.org/10.1109/access.2026.3655392
Cite the original work for its findings. Save a collection to share your selection of sources.