Gpu
50 items across the graph — tagged with Gpu.
From the graph · 50
Tensors and Dynamic neural networks in Python with strong GPU acceleration
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
The fastai deep learning library
Machine Learning Engineering Open Book
Suite of tools for deploying and training deep learning models using the JVM. Highlights include model import for keras, tensorflow, and onnx/pytorch, a modular…
Open3D: A Modern Library for 3D Data Processing
Open Machine Learning Compiler Framework
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faste…
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Line…
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and infer…
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it ins…
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Z…
NVIDIA cuML: GPU-Accelerated Machine Learning
cuML - RAPIDS Machine Learning Library
Time series forecasting with PyTorch
On-device AI across mobile, embedded and edge for PyTorch
One delightful Ruby framework for every major AI provider. Build AI agents, chatbots, RAG apps, and multimodal workflows in beautiful, expressive code.
Achieve state of the art inference performance with modern accelerators on Kubernetes
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwel…
Open-source deep-learning framework for building, training, and fine-tuning deep learning models using state-of-the-art Physics-ML methods
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
A Pythonic framework to simplify AI service building
Resource scheduling and cluster management for AI
General purpose GPU compute framework built on Vulkan to support 1000s of cross vendor graphics cards (AMD, Qualcomm, NVIDIA & friends). Blazing fast, mobile-en…
Deep Learning Server and CLI for Torch and TensorRT
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
A high performance anime upscaler
Ultrafast serverless GPU inference, sandboxes, and background jobs
Simulation of spiking neural networks (SNNs) using PyTorch.
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
🔥 Real-time NVIDIA GPU dashboard
Nvidia GPU exporter for prometheus using nvidia-smi binary
Computations and statistics on manifolds with geometric structures.
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
🌊 Julia software for fast, friendly, flexible, ocean-flavored fluid dynamics on CPUs and GPUs
Extension for Scikit-learn is a seamless way to speed up your Scikit-learn application
Examples of programs built using Modal
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
🎓 系统性大语言模型构建课程|🛠️ 覆盖预训练数据工程、Tokenizer、Transformer、MoE、GPU 编程 (CUDA/Triton)、分布式训练、Scaling Laws、推理优化及对齐 (SFT/RLHF/GRPO)|🚀 6 个渐进式作业 + 代码驱动,建立 LLM 全栈认知体系
RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form bui…
Kubernetes AI Toolchain Operator
JAX in JavaScript – ML library for the web, running on WebGPU & Wasm
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
NVIDIA Merlin is an open source library providing end-to-end GPU-accelerated recommender systems, from feature engineering and preprocessing to training deep le…
BioNeMo Recipes: For building and adapting AI models in drug discovery at scale
