VioletVision-3B
An open vision-language model for captioning and VQA.
Papers96
Modern pretrained vision models achieve strong accuracy but demand substantial GPU memory for fine-tuning, making edge d
paperLIME: Learning Intent-aware Camera Motion from Egocentric VideoAutonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region,
paperChatImage: Navigating Long-Form LLM Answers through Interactive ImagesLarge Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented
paperPerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionWe introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark s
paperU-shaped Multi-granularity Learning for Vision-Language ModelsThe prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lac
paperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQACompositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations
paperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationAccurate breast cancer classification from mammography requires effective integration of complementary information from
paperInhibited Self-Attention: Sharpening Focus in Vision TransformersVision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks. However, their self-attent
Repos5
利用 AI 大模型,一键解说并剪辑视频
repoBlaizzy/mlx-vlmMLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
repopytorch/visionDatasets, Transforms and Models specific to Computer Vision
reposou350121/VLA-Handbook本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战
repovlm-starterA starter kit for training vision-language models.
