Inference
50 items across the graph — tagged with Inference.
From the graph · 50
A high-throughput and memory-efficient inference and serving engine for LLMs
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Cross-platform, customizable ML solutions for live and streaming media.
SGLang is a high-performance serving framework for large language models and multimodal models.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
ncnn is a high-performance neural network inference framework optimized for the mobile platform
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Machine Learning Engineering Open Book
Hugging Face model with 13574 likes. Tags: transformers, safetensors, deepseek_v3, text-generation, conversational, custom_code, arxiv:2501.12948, license:mit,…
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Production ready toolkit to run AI locally
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — a…
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. DoWhy is based on a unified language for causa…
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Uplift modeling and causal inference with machine learning algorithms
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it ins…
Tiny AI for tiny devices
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Z…
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,v…
ALICE (Automated Learning and Intelligence for Causation and Economics) is a Microsoft Research project aimed at applying Artificial Intelligence concepts to ec…
Implement a reasoning LLM in PyTorch from scratch, step by step
Next generation of automated data exploratory analysis and visualization platform.
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with feat…
Hugging Face model with 4171 likes. Tags: transformers, safetensors, deepseek_v3, text-generation, conversational, custom_code, arxiv:2412.19437, eval-results,…
Achieve state of the art inference performance with modern accelerators on Kubernetes
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
Hugging Face model with 3500 likes. Tags: transformers, safetensors, phi, text-generation, nlp, code, en, license:mit, text-generation-inference, endpoints_comp…
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separatio…
📚 Jupyter notebook tutorials for OpenVINO™
Hugging Face model with 3156 likes. Tags: transformers, safetensors, deepseek_v3, text-generation, conversational, custom_code, arxiv:2412.19437, license:mit, e…
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-groun…
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production,…
Community maintained hardware plugin for vLLM on Ascend
Use Hugging Face with JavaScript
Fast ML inference & training for ONNX models in Rust
Turn any computer or edge device into a command center for your computer vision projects.
Open-source inference server and production cluster for all the models your agent needs.
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe not…
Bayesian inference with probabilistic programming.
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
Communicate with an LLM provider using a single interface
