DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong structured visual representations, its lack of native textual alignment hinders its direct application to OVSS. To bridge this gap, we propose DINOde, an ODE-based framework that continuously aligns CLIP text embeddings with the DINO visual manifold. Our approach employs two complementary components: (i) Semantic Text Flow (STF), which evolves text embeddings toward the DINO manifold through a continuous ODE trajectory,
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “DINOde: Continuous Vision-Text Alignment for Open-Vocabulary” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “DINOde: Continuous Vision-Text Alignment for Open-Vocabulary” ≈ “pytorch/vision””
- FuzzySimilar title/name (fuzzy) · 59%microsoft/semantic-kernel →
“Fuzzy title match (0.73): “DINOde: Continuous Vision-Text Alignment for Open-Vocabulary” ≈ “microsoft/semantic-kernel””
- LinkedLinked via arxiv author · 85%Sung-Hoon Yoon →
“DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation”
- LinkedLinked via arxiv author · 85%Hoyong Kwon →
“DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation”
- LinkedLinked via arxiv author · 85%Changgyoon Oh →
“DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation”
- LinkedLinked via arxiv author · 85%Kuk-Jin Yoon →
“DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation”
