Model profile · live graphModel

VioletVision-3B

An open vision-language model for captioning and VQA.

via Hugging Face
1Graph score
101Connections
96Papers
5Repos
0News

Papers96

paperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs

Modern pretrained vision models achieve strong accuracy but demand substantial GPU memory for fine-tuning, making edge d

paperLIME: Learning Intent-aware Camera Motion from Egocentric Video

Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region,

paperChatImage: Navigating Long-Form LLM Answers through Interactive Images

Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented

paperPerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark s

paperU-shaped Multi-granularity Learning for Vision-Language Models

The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lac

paperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations

paperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

Accurate breast cancer classification from mammography requires effective integration of complementary information from

paperInhibited Self-Attention: Sharpening Focus in Vision Transformers

Vision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks. However, their self-attent

Repos5

Topics