Vision
50 items across the graph — tagged with Vision.
From the graph · 50
List of Computer Science courses with video lectures.
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
12 Weeks, 24 Lessons, AI for All!
We write your reusable computer vision tools. 💜
Learn it. Build it. Ship it for others.
Cross-platform, customizable ML solutions for live and streaming media.
Learn OpenCV : C++ and Python Examples
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
YC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ apps
Datasets, Transforms and Models specific to Computer Vision
A toolkit for making real world machine learning and data analysis applications in C++
Open-source simulator for autonomous driving research.
Low-code framework for building custom LLMs, neural networks, and other AI models
🐍 Geometric Computer Vision Library for Spatial AI
Refine high-quality datasets and visual AI models
Fast and Accurate ML in 3 Lines of Code
Effortless data labeling with AI support from Segment Anything and other awesome models.
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-e…
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
Become a cracked AI/ML Research Engineer
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible wi…
Statistical Machine Intelligence & Learning Engine
OpenCV wrapper for .NET
Framework agnostic sliced/tiled inference + interactive ui + error analysis plots
【三年面试五年模拟】AIGC/LLM/AI Agent算法工程师面试秘籍。涵盖AIGC、LLM大模型、AI Agent、具身智能、传统深度学习、自动驾驶、机器学习、计算机视觉、自然语言处理、强化学习、大数据挖掘、世界模型、元宇宙、AGI等AI行业面试笔试干货经验与核心跨周期知识。
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Hugging Face model with 4054 likes. Tags: transformers, safetensors, unlimited-ocr, feature-extraction, baidu, vision-language, ocr, custom_code, image-text-to-…
A python library for self-supervised learning on images.
Superfast AI decision making and intelligent processing of multi-modal data.
Hugging Face model with 3443 likes. Tags: gguf, uncensored, qwen3.6, moe, vision, multimodal, image-text-to-text, en, zh, multilingual
Hugging Face model with 3338 likes. Tags: transformers, safetensors, deepseek_vl_v2, feature-extraction, deepseek, vision-language, ocr, custom_code, image-text…
📚 Jupyter notebook tutorials for OpenVINO™
An open-source cloud-native unified-cloud platform. 开源云原生融合云平台
Hugging Face model with 2743 likes. Tags: transformers, safetensors, locateanything, feature-extraction, nvidia, eagle, vision, object-detection, grounding, arx…
Control Any Computer Using LLMs.
Turn any computer or edge device into a command center for your computer vision projects.
Emgu CV is a cross platform .Net wrapper to the OpenCV image processing library.
UnrealCV: Connecting Computer Vision to Unreal Engine
Collect some World Models for Autonomous Driving (and Robotic, etc.) papers.
Unified multimodal backend for AI data apps
A Unified Semi-Supervised Learning Codebase (NeurIPS'22)
Build computer vision models in a fraction of the time and with less data.
Interactively explore unstructured datasets from your dataframe.
Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.
Videos, notes and experiments to understand deep learning
Fully Open Framework for Democratized Multimodal Training
Microsoft AI for Good Lab — Biodiversity research hub. Open-source AI models, edge devices, and tools for biodiversity monitoring and conservation. Your source…
Semantic photo search for the command line
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Autonomous self-evolving agents. Vision-grounded layered memory and self-written skills for LLM agents that operate your computer.
