repoGitHubTrust 82 · PrimaryPublished 2mo agoLive · 2mo ago
vlm-starter
A starter kit for training vision-language models.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownVision-language pretraining at scale →
- LinkedLinked via unknownTransformer →
- LinkedLinked via unknownPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models →
- LinkedLinked via unknownJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA →
- LinkedLinked via unknownAutomating Potential-based Reward Shaping with Vision Language Model Guidance →
Implements
Related to
Implements (incoming)
paperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change CaptioningpaperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning TaskspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperHarnessing Textual Refusal Directions for Multimodal SafetypaperTone-Conditioned Curriculum Learning for Low-Resource Bantu Speech RecognitionpaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperSearch-based Testing of Vision Language Models for In-Car Scene UnderstandingpaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperShow Me Examples: Inferring Visual Concepts from Image SetspaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatorypaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperLearning to Move Before Learning to Do: Task-Agnostic pretraining for VLAspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperWhen Structured Sparse Autoencoders Learn Consistent Concepts Across ModalitiespaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperTCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language ModelspaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
Related to (incoming)
Covers (incoming)
newsBook Review: Domain-Specific Small Language Models by Guglielmo IozzianewsIEEE Rolls Out Large Language Models Virtual Training Coursenews[D] Looking for Machine Learning / Deep Learning Final Year Project Ideas[D]newsCo-pilot, Not Autopilot: A Practical Method for Using Large Language Models in Interventional Cardiology - EMJnewsyou can just watch a language model think now. i built a way to visualize the words AI doesn’t saynewsI trained a vision-language model to play Snake, and so can you. [P]
Related across the graph
paperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperWhen Structured Sparse Autoencoders Learn Consistent Concepts Across ModalitiespaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Modelsglossary_termTransformerpaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning TaskspaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationnewsI trained a vision-language model to play Snake, and so can you. [P]paperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperHarnessing Textual Refusal Directions for Multimodal SafetypaperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatorypaperTCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language ModelspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigationnewsyou can just watch a language model think now. i built a way to visualize the words AI doesn’t saypaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelsmodelVioletVision-3BpaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextnewsBook Review: Domain-Specific Small Language Models by Guglielmo IozziapaperLearning to Move Before Learning to Do: Task-Agnostic pretraining for VLAspaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioningnews[D] Looking for Machine Learning / Deep Learning Final Year Project Ideas[D]newsIEEE Rolls Out Large Language Models Virtual Training CoursepaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperVision-language pretraining at scalepaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperTone-Conditioned Curriculum Learning for Low-Resource Bantu Speech RecognitionpaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperShow Me Examples: Inferring Visual Concepts from Image SetspaperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelsnewsCo-pilot, Not Autopilot: A Practical Method for Using Large Language Models in Interventional Cardiology - EMJpaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionpaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperSearch-based Testing of Vision Language Models for In-Car Scene Understanding
