VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval
Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-language and reference-based approaches have significantly advanced this field, they often rely on text prompts, limited visual examples, or expensive feature matching procedures, making them difficult to scale to large and continuously expanding object repositories. In this work, we propose VocaDet, a sample-driven open-vocabulary object detection and segmentation framework that learns object concepts directly from user-provided positive and negative sample
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%VioletVision-3B →
- PossiblePossibly related (embedding) · 51%VCG-team/DiffSegmenter →
- PossiblePossibly related (embedding) · 48%llmsresearch/llm-flashcards →
- PossiblePossibly related (embedding) · 48%vlm-starter →
- PossiblePossibly related (embedding) · 48%NVIDIA-ISAAC-ROS/isaac_ros_object_detection →
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%ZhiXin Sun →
“VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Ret”
