modelHugging FaceTrust 88 · LabPublished 3mo agoLive · 1mo ago1 graph score
VioletVision-3B
An open vision-language model for captioning and VQA.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownvlm-starter →
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “VioletVision-3B” ≈ “pytorch/vision””
- LinkedLinked via unknownViQ: Text-Aligned Visual Quantized Representations at Any Resolution →
- LinkedLinked via unknownPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models →
- LinkedLinked via unknownJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA →
- LinkedLinked via unknownHarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models →
Related to
Has model (incoming)
paperViQ: Text-Aligned Visual Quantized Representations at Any ResolutionpaperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperHarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal ModelspaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change CaptioningpaperPerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionpaperZero-Gated Language-conditioned Human Motion PredictionpaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperLow-cost concept-based localized explanations: How far can we get with training-free approaches?paperLatent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language ModelspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperBefore Thinking, Learn to Decide: Proactive Routing for Efficient Visual ReasoningpaperReweighting Framewise Attention in Video Transformers for Facial Expression UnderstandingpaperOpen-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D DetectorspaperHarnessing Textual Refusal Directions for Multimodal SafetypaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM InferencepaperERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMspaperVisual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?paperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperPerceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningpaperFurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action ModelpaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperShow Me Examples: Inferring Visual Concepts from Image SetspaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperSeek to Segment: Active Perception for Panoramic Referring SegmentationpaperLIME: Learning Intent-aware Camera Motion from Egocentric VideopaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperReasoning LLM Improves Speaker Recognition in Long-form TV DramaspaperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperVision as Unified Multimodal GenerationpaperCognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and EditingpaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperStoryTeller: Training-Free Narrative Grounding for Long-Form Audio DescriptionpaperEvidence-Backed Video Question AnsweringpaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperVision-language pretraining at scalepaperFine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT UnderstandingpaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUspaperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperVisual Access Boundaries in Vision-Language Model ReasoningpaperVisually Grounded Self-Reflection for Vision-Language Models via Reinforcement LearningpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperDART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language NavigationpaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperBrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagepaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperAirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelspaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperU-shaped Multi-granularity Learning for Vision-Language ModelspaperParameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningpaperSceneBind: Binding What and Where Across Vision, Audio and LanguagepaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperVision-Language Assistant for Emotional Reactions to Risky DrivingpaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperSearch-based Testing of Vision Language Models for In-Car Scene UnderstandingpaperSearching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language ModelspaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperAnticipate Before Acting: Future-State-Conditioned Vision-Language NavigationpaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperRCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile GeneralizationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperContext-structured Video Anomaly Detection with Large Vision-Language ModelspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperTest-Time Training for Modality Order Consistency in Vision-Language ModelspaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperDINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Related to (incoming)
Related across the graph
paperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUspaperLIME: Learning Intent-aware Camera Motion from Egocentric VideopaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperPerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionpaperU-shaped Multi-granularity Learning for Vision-Language ModelspaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationpaperInhibited Self-Attention: Sharpening Focus in Vision Transformersrepolinyqh/NarratoAIpaperVision-Language Assistant for Emotional Reactions to Risky DrivingpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperAnticipate Before Acting: Future-State-Conditioned Vision-Language NavigationpaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperBefore Thinking, Learn to Decide: Proactive Routing for Efficient Visual ReasoningpaperVisual Access Boundaries in Vision-Language Model ReasoningpaperVisually Grounded Self-Reflection for Vision-Language Models via Reinforcement LearningpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperDART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language NavigationpaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperTest-Time Training for Modality Order Consistency in Vision-Language ModelspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMspaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperParameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningpaperFine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT UnderstandingrepoBlaizzy/mlx-vlmpaperAirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelspaperOpen-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D DetectorspaperSeek to Segment: Active Perception for Panoramic Referring SegmentationpaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperHarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal ModelspaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperContext-structured Video Anomaly Detection with Large Vision-Language ModelspaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperSceneBind: Binding What and Where Across Vision, Audio and LanguagepaperRCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile GeneralizationpaperTORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperHarnessing Textual Refusal Directions for Multimodal SafetypaperZero-Gated Language-conditioned Human Motion Predictionrepopytorch/visionpaperEvidence-Backed Video Question AnsweringpaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperReweighting Framewise Attention in Video Transformers for Facial Expression UnderstandingpaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperVisual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?paperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperLatent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language ModelspaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM InferencepaperDINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic SegmentationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperSearching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language ModelspaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperPerceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningpaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEnhancing Part-Level Point Grounding for Any Open-Source MLLMspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperRSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change CaptioningpaperViQ: Text-Aligned Visual Quantized Representations at Any ResolutionpaperBrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagepaperVision-language pretraining at scalepaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperLow-cost concept-based localized explanations: How far can we get with training-free approaches?paperCognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and EditingpaperEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelspaperThe Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMspaperPaying More Attention to Visual Tokens in Self-Evolving Large Multimodal ModelspaperReasoning LLM Improves Speaker Recognition in Long-form TV Dramasreposou350121/VLA-HandbookpaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperAsk, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency RewardspaperVision as Unified Multimodal GenerationpaperShow Me Examples: Inferring Visual Concepts from Image SetspaperSeeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialoguepaperAnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language ModelspaperFurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Modelrepovlm-starterpaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM InferencepaperStoryTeller: Training-Free Narrative Grounding for Long-Form Audio DescriptionpaperSearch-based Testing of Vision Language Models for In-Car Scene Understanding
