arXiv · 2410.16028
Few-shot target-driven instance detection based on open-vocabulary object detection models
Abstract
Current large open vision models could be useful for one and few-shot object recognition. Nevertheless, gradient-based re-training solutions are costly. On the other hand, open-vocabulary object detection models bring closer visual and textual concepts in the same latent space, allowing zero-shot detection via prompting at small computational cost. We propose a lightweight method to turn the latter into a one-shot or few-shot object recognition models without requiring textual descriptions. Our experiments on the TEgO dataset using the YOLO-World model as a base show that performance increases with the model size, the number of examples and the use of image augmentation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ben Crulis, Barthelemy Serres, Cyril De Runz, Gilles Venturini. 2024-10-21. Few-shot target-driven instance detection based on open-vocabulary object detection models. https://arxiv.org/abs/2410.16028
Cite the original work for its findings. Save a collection to share your selection of sources.