Vllm
29 items across the graph — tagged with Vllm.
From the graph · 29
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Hugging Face model with 5106 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…
Hugging Face model with 4919 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…
Hugging Face model with 4720 likes. Tags: vllm, safetensors, mixtral, fr, it, de, es, en, base_model:mistralai/Mixtral-8x7B-v0.1, base_model:finetune:mistralai/…
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
Hugging Face model with 2705 likes. Tags: vllm, safetensors, mistral, mistral-common, base_model:mistralai/Mistral-7B-v0.3, base_model:finetune:mistralai/Mistra…
Community maintained hardware plugin for vLLM on Ascend
The Runpod worker template for serving our large language model endpoints. Powered by vLLM.
Python package for LLM compression
Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
A PyTorch native library for training speculative decoding models
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference
The open-source AI platform for enterprises that can't send data to the cloud. OpenAI-compatible API, full management dashboard, zero data egress.
Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
OpenAI Router 轻量级、持久化、零配置的 OpenAI API 统一网关
Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and com…
Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed2…
Dependency/tool package detected from repository manifests (vllm.general_plugins).
Dependency/tool package detected from repository manifests (vllm).
