Topic

Cuda

50 items across the graph — tagged with Cuda.

From the graph · 50

repo
vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

repo
sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

repo
tracel-ai/burn

Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.

repo
isl-org/Open3D

Open3D: A Modern Library for 3D Data Processing

repo
LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

repo
replicate/cog

Containers for machine learning

repo
catboost/catboost

A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks fo…

repo
InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

repo
gpustack/gpustack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

repo
NVIDIA/cuml

NVIDIA cuML: GPU-Accelerated Machine Learning

repo
rapidsai/cuml

cuML - RAPIDS Machine Learning Library

repo
NVIDIAGameWorks/kaolin

A PyTorch Library for Accelerating 3D Deep Learning Research

repo
iree-org/iree

A retargetable MLIR-based machine learning compiler and runtime toolkit.

repo
NVIDIA/TransformerEngine

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwel…

repo
BBuf/how-to-optim-algorithm-in-cuda

how to optimize some algorithm in cuda.

repo
pytorch/TensorRT

PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT

repo
containers/ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production,…

repo
NVIDIA/skills

Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows en…

repo
withcatai/node-llama-cpp

Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level

repo
dphnAI/sonar

Large-scale LLM inference engine

repo
dphnAI/aphrodite-engine

Large-scale LLM inference engine

repo
beam-cloud/beta9

Ultrafast serverless GPU inference, sandboxes, and background jobs

repo
tenstorrent/tt-metal

:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.

repo
psalias2006/gpu-hot

🔥 Real-time NVIDIA GPU dashboard

repo
uccl-project/uccl

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

repo
SemiAnalysisAI/InferenceX

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™…

repo
NVIDIA/raft

RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form bui…

repo
insight-platform/Savant

Python Computer Vision & Video Analytics Framework With Batteries Included

repo
NVIDIA/cuvs

cuVS - a library for vector search and clustering on the GPU

repo
brucefan1983/GPUMD

Graphics Processing Units Molecular Dynamics

repo
jmaczan/tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

repo
invergent-ai/surogate

Training/Fine-tuning at the speed of light

repo
MarioSieg/magnetron

A zero-dependency ML framework in C with a modern Python API for full control over execution and memory.

repo
pegainfer-project/pegainfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

repo
openinfer-project/openinfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

repo
hybridgroup/yzma

Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.

repo
AmusementClub/vs-mlrt

Efficient CPU/GPU ML Runtimes for VapourSynth (with built-in support for waifu2x, DPIR, RealESRGANv2/v3, Real-CUGAN, RIFE, SCUNet, ArtCNN and more!)

repo
Entrpi/ds4-on-spark

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous…

repo
avifenesh/memra

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned de…

repo
ModelEngine-Group/unified-cache-management

Persist and reuse KV Cache to speedup your LLM.

repo
raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

repo
ot-triton-lab/flash-sinkhorn

The official repository of FlashSinkhorn [ICML 2026 Oral]

repo
giannisanni/pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-…

repo
kibae/onnxruntime-server

ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.

repo
traceopt-ai/traceml

Open-source performance diagnostics for PyTorch training runs.

repo
inclusionAI/Awex

A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from training to inference in RL wo…

repo
NVIDIA/nvImageCodec

A nvImageCodec library of GPU- and CPU- accelerated codecs featuring a unified interface

repo
mlx-node/mlx-node
repo
gammahazard/locate-anything

Sleek, mobile-friendly web UI for NVIDIA LocateAnything-3B — open-vocabulary object detection & grounding on your own GPU, via one docker compose up.

repo
jdermody/brightwire

Bright Wire is an open source machine learning library for .NET with GPU support (via CUDA)

Related topics