Inference Server
11 items across the graph — tagged with Inference Server.
From the graph · 11
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production,…
Turn any computer or edge device into a command center for your computer vision projects.
Open-source inference server and production cluster for all the models your agent needs.
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, a…
The simplest way to serve AI/ML models in production
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and im…
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and im…
