Llm Inference
50 items across the graph — tagged with Llm Inference.
From the graph · 50
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Z…
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-groun…
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe not…
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that…
LLMs as Copilots for Theorem Proving in Lean
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)
“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space
Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )
RDNA-native LLM inference engine in Rust.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs.
RDNA-native LLM inference engine in Rust.
Explore the unknown, build the future, own your data.
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-ho…
Everything you need to know about LLM inference
🌱 EcoLogits tracks the energy consumption and environmental footprint of using generative AI models through APIs.
电子鹦鹉 / Toy Language Model
Run generative AI models in sophgo BM1684X/BM1688
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Minimal yet performant LLM examples in pure JAX
Aggregates compute from spare GPU capacity
A light llama-like llm inference framework based on the triton kernel.
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…
Pivotal Token Search
Open-source local AI workspace — advancing on-device inference.
Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source
lcpp is a dart implementation of llama.cpp used by the mobile artificial intelligence distribution (maid)
A Lightweight LLM Inference Performance Simulator
User friendly GUI for configuring and launching llama.cpp
Get more intelligence from every bit. Better quantization formats and smarter calibration let larger, stronger models run smoothly on the hardware you already o…
Telegram bot for different language models. Supports system prompts and images
Neutral, reproducible benchmark for local LLMs on Apple Silicon (Mac · iPhone · iPad) — MLX, llama.cpp, CoreML, Apple Foundation Models
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process via FFI. Chat, streaming, embeddings, RAG and vect…
FreeGPTHub: A truly free, unified GPT API gateway—unlike “fake-free” alternatives that hide paywalls behind quotas, trials, or mandatory top-ups. (真正免费的GPT统一接口,…
Browser based Interface for Generative AI. Chat/Agent/Taskmanager Hybrid.
