Cost-efficient Active Learning for Referring Image Segmentation and Grounding
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually similar ones. We tackle this by formulating active learning (AL) for VG under the realistic setting where only raw images are available without accompanying text. Since ground-truth text is unavailable, sample selection must estimate which images contain ambiguous regions that would require discriminative referring expressions. To address this, we generate auxi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo →
“Fuzzy title match (0.73): “Cost-efficient Active Learning for Referring Image Segmentat” ≈ “Tongyi-MAI/Z-Image-Turbo””
- LinkedLinked via arxiv author · 85%Junbeom Hong →
“Cost-efficient Active Learning for Referring Image Segmentation and Grounding”
- LinkedLinked via arxiv author · 85%Seonghoon Yu →
“Cost-efficient Active Learning for Referring Image Segmentation and Grounding”
- LinkedLinked via arxiv author · 85%Hyung Rok Jung →
“Cost-efficient Active Learning for Referring Image Segmentation and Grounding”
- LinkedLinked via arxiv author · 85%Sundong Kim →
“Cost-efficient Active Learning for Referring Image Segmentation and Grounding”
- LinkedLinked via arxiv author · 85%Jeany Son →
“Cost-efficient Active Learning for Referring Image Segmentation and Grounding”
- FuzzySimilar title/name (fuzzy) · 84%amitness/learning →
“Fuzzy title match (0.92): “Cost-efficient Active Learning for Referring Image Segmentat” ≈ “amitness/learning””
