Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /vlm-starter
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 2mo agoLive · 2mo ago

vlm-starter

A starter kit for training vision-language models.

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via unknownVision-language pretraining at scale →
  • LinkedLinked via unknownTransformer →
  • LinkedLinked via unknownPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models →
  • LinkedLinked via unknownAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards →
  • LinkedLinked via unknownJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA →
  • LinkedLinked via unknownTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference →
  • LinkedLinked via unknownLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline) →
  • LinkedLinked via unknownAutomating Potential-based Reward Shaping with Vision Language Model Guidance →

Implements

paperVision-language pretraining at scale

Related to

glossary_termTransformer

Implements (incoming)

paperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change CaptioningpaperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning TaskspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperHarnessing Textual Refusal Directions for Multimodal SafetypaperTone-Conditioned Curriculum Learning for Low-Resource Bantu Speech RecognitionpaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperSearch-based Testing of Vision Language Models for In-Car Scene UnderstandingpaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperShow Me Examples: Inferring Visual Concepts from Image SetspaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatorypaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperLearning to Move Before Learning to Do: Task-Agnostic pretraining for VLAspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperWhen Structured Sparse Autoencoders Learn Consistent Concepts Across ModalitiespaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperTCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language ModelspaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Related to (incoming)

modelVioletVision-3B

Covers (incoming)

newsBook Review: Domain-Specific Small Language Models by Guglielmo IozzianewsIEEE Rolls Out Large Language Models Virtual Training Coursenews[D] Looking for Machine Learning / Deep Learning Final Year Project Ideas[D]newsCo-pilot, Not Autopilot: A Practical Method for Using Large Language Models in Interventional Cardiology - EMJnewsyou can just watch a language model think now. i built a way to visualize the words AI doesn’t saynewsI trained a vision-language model to play Snake, and so can you. [P]

Related across the graph

paperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperWhen Structured Sparse Autoencoders Learn Consistent Concepts Across ModalitiespaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Modelsglossary_termTransformerpaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning TaskspaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationnewsI trained a vision-language model to play Snake, and so can you. [P]paperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperHarnessing Textual Refusal Directions for Multimodal SafetypaperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatorypaperTCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language ModelspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigationnewsyou can just watch a language model think now. i built a way to visualize the words AI doesn’t saypaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelsmodelVioletVision-3BpaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextnewsBook Review: Domain-Specific Small Language Models by Guglielmo IozziapaperLearning to Move Before Learning to Do: Task-Agnostic pretraining for VLAspaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioningnews[D] Looking for Machine Learning / Deep Learning Final Year Project Ideas[D]newsIEEE Rolls Out Large Language Models Virtual Training CoursepaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperVision-language pretraining at scalepaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperTone-Conditioned Curriculum Learning for Low-Resource Bantu Speech RecognitionpaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperShow Me Examples: Inferring Visual Concepts from Image SetspaperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelsnewsCo-pilot, Not Autopilot: A Practical Method for Using Large Language Models in Interventional Cardiology - EMJpaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionpaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperSearch-based Testing of Vision Language Models for In-Car Scene Understanding
Knowledge path·PHoloCount: A Holistic Visual Counting Benchmark for MLLMs→PDynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents→PWhen Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities→Rvlm-starter

Topics

visionmultimodal

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Graph trust82Primary
Graph score1