repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
Somnusochi/VLM-AutoYOLO
AI Auto Annotation & YOLO Training Pipeline, End-to-end object detection auto-labeling and YOLO training platform. VLM-powered annotation with NVIDIA LocateAnything-3B, manual refinement, one-click YOLO training, video keyframe extraction, and model validation. Supports image and video.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning →
- PossiblePossibly related (embedding) · 54%AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models →
- PossiblePossibly related (embedding) · 53%Object-centric LeJEPA →
- PossiblePossibly related (embedding) · 48%OctoSense: Self-Supervised Learning for Multimodal Robot Perception →
- PossiblePossibly related (embedding) · 48%AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting →
- PossiblePossibly related (embedding) · 55%NVIDIA AI Introduces ASPIRE: A Self-Improving Robotics Framework Reaching 31% Zero-Shot on LIBERO-Pro Long Tasks - MarkTechPost →
- PossiblePossibly related (embedding) · 47%G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement →
- PossiblePossibly related (embedding) · 51%VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection →
Covers
Implements
paperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperObject-centric LeJEPApaperOctoSense: Self-Supervised Learning for Multimodal Robot PerceptionpaperAdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting
Covers (incoming)
Implements (incoming)
paperG2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementpaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionpaperWhareformer: Learning to Track What is Where in Long Egocentric VideospaperSAM-MT: Real-Time Interactive Multi-Target Video Segmentation
Related across the graph
paperG2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementpaperObject-centric LeJEPAnewsInto the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-TuningpaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionnewsBuild a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills | NVIDIA Technical Blog - NVIDIA DevelopernewsNVIDIA AI Introduces ASPIRE: A Self-Improving Robotics Framework Reaching 31% Zero-Shot on LIBERO-Pro Long Tasks - MarkTechPostpaperSAM-MT: Real-Time Interactive Multi-Target Video SegmentationpaperAdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object CountingpaperWhareformer: Learning to Track What is Where in Long Egocentric VideospaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperOctoSense: Self-Supervised Learning for Multimodal Robot Perception
