repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 22d ago
pytorch/vision
Datasets, Transforms and Models specific to Computer Vision
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers →
- PossiblePossibly related (embedding) · 48%[D] Looking for Machine Learning / Deep Learning Final Year Project Ideas[D] →
- PossiblePossibly related (embedding) · 48%Transformer Geometry Observatory TGO-II: Representational Similarity Observatory →
- PossiblePossibly related (embedding) · 48%Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning →
- FuzzySimilar title/name (fuzzy) · 84%Vision-language pretraining at scale →
“Fuzzy title match (0.92): “Vision-language pretraining at scale” ≈ “pytorch/vision””
- FuzzySimilar title/name (fuzzy) · 84%Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding →
“Fuzzy title match (0.92): “Fine-Grained Vision-Language Pretraining with Organ-Conditio” ≈ “pytorch/vision””
- FuzzySimilar title/name (fuzzy) · 84%MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation →
“Fuzzy title match (0.92): “MonoIR-RS: Infrared Remote Sensing Vision-Language Learning ” ≈ “pytorch/vision””
- FuzzySimilar title/name (fuzzy) · 84%FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers →
“Fuzzy title match (0.92): “FlexViT: A Flexible FPGA-based Accelerator for Edge Vision T” ≈ “pytorch/vision””
Implements
paperDPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View TransformerspaperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatorypaperVision-language pretraining at scalepaperFine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT UnderstandingpaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUspaperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationpaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperVisual Access Boundaries in Vision-Language Model ReasoningpaperVisually Grounded Self-Reflection for Vision-Language Models via Reinforcement LearningpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperDART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language NavigationpaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperBrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagepaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperAirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelspaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperU-shaped Multi-granularity Learning for Vision-Language ModelspaperParameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningpaperSceneBind: Binding What and Where Across Vision, Audio and LanguagepaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelspaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperVision-Language Assistant for Emotional Reactions to Risky DrivingpaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperSearch-based Testing of Vision Language Models for In-Car Scene UnderstandingpaperSearching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language ModelspaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperAnticipate Before Acting: Future-State-Conditioned Vision-Language NavigationpaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperRCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile GeneralizationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperContext-structured Video Anomaly Detection with Large Vision-Language ModelspaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperTest-Time Training for Modality Order Consistency in Vision-Language ModelspaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperDINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic SegmentationpaperGS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model PriorspaperUniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action ModelspaperRollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-TrainingpaperStyle or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision EmbeddingspaperSeeing Red, Thinking Bad: Color Bias in Vision Language ModelspaperOn the Robustness of Temporal Vision-Language Models for Surgical Endoscopy VideospaperViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models
Covers
Implements (incoming)
Related to (incoming)
Related across the graph
paperEfficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUspaperU-shaped Multi-granularity Learning for Vision-Language ModelspaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperToken-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer ClassificationpaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperVision-Language Assistant for Emotional Reactions to Risky DrivingpaperSeeing Red, Thinking Bad: Color Bias in Vision Language ModelspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperAnticipate Before Acting: Future-State-Conditioned Vision-Language NavigationpaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationnewsInto the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-TuningpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperVisual Access Boundaries in Vision-Language Model ReasoningpaperVisually Grounded Self-Reflection for Vision-Language Models via Reinforcement LearningpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperOn the Robustness of Temporal Vision-Language Models for Surgical Endoscopy VideospaperDART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language NavigationpaperALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertspaperTest-Time Training for Modality Order Consistency in Vision-Language ModelspaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperAUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingpaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperUniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action ModelspaperParameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningpaperFine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT UnderstandingpaperViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation ModelspaperAirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelspaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperContext-structured Video Anomaly Detection with Large Vision-Language ModelspaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperSceneBind: Binding What and Where Across Vision, Audio and LanguagepaperRCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile GeneralizationpaperTraining Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperStyle or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision EmbeddingspaperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatorypaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperGenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language ModelspaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperDINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic SegmentationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperSearching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language ModelspaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperVision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language ModelsmodelVioletVision-3BpaperMulti-Resolution Feature Stem for Diabetic Retinopathy lesion segmentationpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperDeform360: A Massive Multi-view Visuotactile Dataset for Deformable World ModelspaperGS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model PriorspaperRollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-TrainingpaperWhareformer: Learning to Track What is Where in Long Egocentric Videosnews[D] Looking for Machine Learning / Deep Learning Final Year Project Ideas[D]paperBrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagepaperVision-language pretraining at scalepaperDPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View TransformerspaperSearch-based Testing of Vision Language Models for In-Car Scene Understanding
