Cuda
50 items across the graph — tagged with Cuda.
From the graph · 50
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
Open3D: A Modern Library for 3D Data Processing
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Containers for machine learning
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks fo…
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
NVIDIA cuML: GPU-Accelerated Machine Learning
cuML - RAPIDS Machine Learning Library
A PyTorch Library for Accelerating 3D Deep Learning Research
A retargetable MLIR-based machine learning compiler and runtime toolkit.
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwel…
how to optimize some algorithm in cuda.
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production,…
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows en…
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
Large-scale LLM inference engine
Large-scale LLM inference engine
Ultrafast serverless GPU inference, sandboxes, and background jobs
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
🔥 Real-time NVIDIA GPU dashboard
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™…
RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form bui…
Python Computer Vision & Video Analytics Framework With Batteries Included
cuVS - a library for vector search and clustering on the GPU
Graphics Processing Units Molecular Dynamics
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Training/Fine-tuning at the speed of light
A zero-dependency ML framework in C with a modern Python API for full control over execution and memory.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Efficient CPU/GPU ML Runtimes for VapourSynth (with built-in support for waifu2x, DPIR, RealESRGANv2/v3, Real-CUGAN, RIFE, SCUNet, ArtCNN and more!)
Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous…
Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned de…
Persist and reuse KV Cache to speedup your LLM.
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
The official repository of FlashSinkhorn [ICML 2026 Oral]
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-…
ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.
Open-source performance diagnostics for PyTorch training runs.
A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from training to inference in RL wo…
A nvImageCodec library of GPU- and CPU- accelerated codecs featuring a unified interface
Sleek, mobile-friendly web UI for NVIDIA LocateAnything-3B — open-vocabulary object detection & grounding on your own GPU, via one docker compose up.
Bright Wire is an open source machine learning library for .NET with GPU support (via CUDA)
