Topic

Vllm

29 items across the graph — tagged with Vllm.

From the graph · 29

repo
LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

repo
kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

model
openai/gpt-oss-120b

Hugging Face model with 5106 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…

model
openai/gpt-oss-20b

Hugging Face model with 4919 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…

model
mistralai/Mixtral-8x7B-Instruct-v0.1

Hugging Face model with 4720 likes. Tags: vllm, safetensors, mixtral, fr, it, de, es, en, base_model:mistralai/Mixtral-8x7B-v0.1, base_model:finetune:mistralai/…

repo
skyzh/tiny-llm

A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.

repo
PaddlePaddle/FastDeploy

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

model
mistralai/Mistral-7B-Instruct-v0.3

Hugging Face model with 2705 likes. Tags: vllm, safetensors, mistral, mistral-common, base_model:mistralai/Mistral-7B-v0.3, base_model:finetune:mistralai/Mistra…

repo
vllm-project/vllm-ascend

Community maintained hardware plugin for vLLM on Ascend

repo
runpod-workers/worker-vllm

The Runpod worker template for serving our large language model endpoints. Powered by vLLM.

repo
FujitsuResearch/OneCompression

Python package for LLM compression

repo
guoqingbao/xinfer

Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.

repo
mudler/vllm.cpp

a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features

repo
AstraNetLab/CacheRoute

CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…

repo
matrixhub-ai/matrixhub

An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.

repo
lightseekorg/TorchSpec

A PyTorch native library for training speculative decoding models

repo
novitalabs/pegaflow

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

repo
BJTU-ANT/CacheRoute

CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…

repo
ai-runway/airunway

✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers

repo
kaito-project/airunway

✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers

repo
xcena-dev/maru

High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference

repo
xinity-ai/xinity-ai

The open-source AI platform for enterprises that can't send data to the cloud. OpenAI-compatible API, full management dashboard, zero data egress.

repo
outsourc-e/bench-loop

Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.

repo
timothystewart6/vllm-gb10

Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).

repo
shell-nlp/openai_router

OpenAI Router 轻量级、持久化、零配置的 OpenAI API 统一网关

repo
slb350/open-agent-sdk-rust

Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and com…

repo
AEON-7/Aeon-Bench-Pod

Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed2…

tool
vllm.general_plugins

Dependency/tool package detected from repository manifests (vllm.general_plugins).

tool
vllm

Dependency/tool package detected from repository manifests (vllm).

Related topics