Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Models
  3. /VioletVision-3B
Read original ↗
modelHugging FaceTrust 88 · LabPublished 3mo agoLive · 1mo ago1 graph score

VioletVision-3B

An open vision-language model for captioning and VQA.

visionmultimodal

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via unknownvlm-starter →
  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →

    “Fuzzy title match (0.92): “VioletVision-3B” ≈ “pytorch/vision””

  • LinkedLinked via unknownViQ: Text-Aligned Visual Quantized Representations at Any Resolution →
  • LinkedLinked via unknownPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models →
  • LinkedLinked via unknownAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards →
  • LinkedLinked via unknownJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA →
  • LinkedLinked via unknownHarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models →
  • LinkedLinked via unknownTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference →

Related to

repovlm-starterrepopytorch/vision

Has model (incoming)

paperViQ: Text-Aligned Visual Quantized Representations at Any ResolutionpaperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperHarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal ModelspaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change CaptioningpaperPerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionpaperZero-Gated Language-conditioned Human Motion PredictionpaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperLow-cost concept-based localized explanations: How far can we get with training-free approaches?paperLatent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language ModelspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperBefore Thinking, Learn to Decide: Proactive Routing for Efficient Visual ReasoningpaperReweighting Framewise Attention in Video Transformers for Facial Expression UnderstandingpaperOpen-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D DetectorspaperHarnessing Textual Refusal Directions for Multimodal SafetypaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM InferencepaperERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMspaperVisual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?paperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperPerceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningpaperFurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action ModelpaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperShow Me Examples: Inferring Visual Concepts from Image SetspaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperSeek to Segment: Active Perception for Panoramic Referring SegmentationpaperLIME: Learning Intent-aware Camera Motion from Egocentric VideopaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperReasoning LLM Improves Speaker Recognition in Long-form TV DramaspaperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperVision as Unified Multimodal GenerationpaperCognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and EditingpaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperStoryTeller: Training-Free Narrative Grounding for Long-Form Audio DescriptionpaperEvidence-Backed Video Question AnsweringpaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperVision-language pretraining at scalepaperFine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT UnderstandingpaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUspaperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperVisual Access Boundaries in Vision-Language Model ReasoningpaperVisually Grounded Self-Reflection for Vision-Language Models via Reinforcement LearningpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperDART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language NavigationpaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperBrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagepaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperAirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelspaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperU-shaped Multi-granularity Learning for Vision-Language ModelspaperParameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningpaperSceneBind: Binding What and Where Across Vision, Audio and LanguagepaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperVision-Language Assistant for Emotional Reactions to Risky DrivingpaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperSearch-based Testing of Vision Language Models for In-Car Scene UnderstandingpaperSearching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language ModelspaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperAnticipate Before Acting: Future-State-Conditioned Vision-Language NavigationpaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperRCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile GeneralizationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperContext-structured Video Anomaly Detection with Large Vision-Language ModelspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperTest-Time Training for Modality Order Consistency in Vision-Language ModelspaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperDINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

Related to (incoming)

reposou350121/VLA-HandbookrepoBlaizzy/mlx-vlmrepolinyqh/NarratoAI

Related across the graph

paperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUspaperLIME: Learning Intent-aware Camera Motion from Egocentric VideopaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperPerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionpaperU-shaped Multi-granularity Learning for Vision-Language ModelspaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationpaperInhibited Self-Attention: Sharpening Focus in Vision Transformersrepolinyqh/NarratoAIpaperVision-Language Assistant for Emotional Reactions to Risky DrivingpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperAnticipate Before Acting: Future-State-Conditioned Vision-Language NavigationpaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperBefore Thinking, Learn to Decide: Proactive Routing for Efficient Visual ReasoningpaperVisual Access Boundaries in Vision-Language Model ReasoningpaperVisually Grounded Self-Reflection for Vision-Language Models via Reinforcement LearningpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperDART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language NavigationpaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperTest-Time Training for Modality Order Consistency in Vision-Language ModelspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMspaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperParameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningpaperFine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT UnderstandingrepoBlaizzy/mlx-vlmpaperAirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelspaperOpen-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D DetectorspaperSeek to Segment: Active Perception for Panoramic Referring SegmentationpaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperHarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal ModelspaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperContext-structured Video Anomaly Detection with Large Vision-Language ModelspaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperSceneBind: Binding What and Where Across Vision, Audio and LanguagepaperRCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile GeneralizationpaperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperHarnessing Textual Refusal Directions for Multimodal SafetypaperZero-Gated Language-conditioned Human Motion Predictionrepopytorch/visionpaperEvidence-Backed Video Question AnsweringpaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperReweighting Framewise Attention in Video Transformers for Facial Expression UnderstandingpaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperVisual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?paperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperLatent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language ModelspaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM InferencepaperDINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic SegmentationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperSearching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language ModelspaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperPerceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningpaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change CaptioningpaperViQ: Text-Aligned Visual Quantized Representations at Any ResolutionpaperBrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagepaperVision-language pretraining at scalepaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperLow-cost concept-based localized explanations: How far can we get with training-free approaches?paperCognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and EditingpaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperReasoning LLM Improves Speaker Recognition in Long-form TV Dramasreposou350121/VLA-HandbookpaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperVision as Unified Multimodal GenerationpaperShow Me Examples: Inferring Visual Concepts from Image SetspaperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperFurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Modelrepovlm-starterpaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperStoryTeller: Training-Free Narrative Grounding for Long-Form Audio DescriptionpaperSearch-based Testing of Vision Language Models for In-Car Scene Understanding
Knowledge path·PEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs→PLIME: Learning Intent-aware Camera Motion from Egocentric Video→PChatImage: Navigating Long-Form LLM Answers through Interactive Images→MVioletVision-3B

Topics

visionmultimodal
View full model profile →

Explore

Search similar →Knowledge graph →All models →Full intelligence feed →
Graph trust88Lab
Graph score1