Topic

Llm Inference

50 items across the graph — tagged with Llm Inference.

From the graph · 50

repo
ray-project/ray

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

repo
liguodongiot/llm-action

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

repo
Lightning-AI/litgpt

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

repo
bentoml/BentoML

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

repo
InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

repo
drumih/turbo-fieldfare

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

repo
cactus-compute/cactus

Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.

repo
gpustack/gpustack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

repo
lemonade-sdk/lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Z…

repo
spiceai/spiceai

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-groun…

repo
b4rtaz/distributed-llama

Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

repo
Nano-Collective/nanocoder

An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe not…

repo
neuron-core/neuron-ai

The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that…

repo
lean-dojo/LeanCopilot

LLMs as Copilots for Theorem Proving in Lean

repo
ovg-project/kvcached

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

repo
foldl/chatllm.cpp

Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)

repo
tingaicompass/AI-Compass

“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。

repo
eastriverlee/LLM.swift

LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.

repo
jmaczan/tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

repo
felladrin/MiniSearch

Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space

repo
NotPunchnox/rkllama

Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )

repo
warpfront/hipfire

RDNA-native LLM inference engine in Rust.

repo
ome-projects/ome

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

repo
expectedparrot/edsl

Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs.

repo
Kaden-Schutt/hipfire

RDNA-native LLM inference engine in Rust.

repo
NPC-Worldwide/incognide

Explore the unknown, build the future, own your data.

repo
jaylfc/taOS

Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-ho…

repo
bentoml/llm-inference-handbook

Everything you need to know about LLM inference

repo
mlco2/ecologits

🌱 EcoLogits tracks the energy consumption and environmental footprint of using generative AI models through APIs.

repo
bd4sur/Nano

电子鹦鹉 / Toy Language Model

repo
sophgo/LLM-TPU

Run generative AI models in sophgo BM1684X/BM1688

repo
AstraNetLab/CacheRoute

CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…

repo
matrixhub-ai/matrixhub

An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.

repo
jax-ml/jax-llm-examples

Minimal yet performant LLM examples in pure JAX

repo
numindai/nuextract
repo
kalavai-net/kalavai-client

Aggregates compute from spare GPU capacity

repo
harleyszhang/lite_llama

A light llama-like llm inference framework based on the triton kernel.

repo
BJTU-ANT/CacheRoute

CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…

repo
codelion/pts

Pivotal Token Search

repo
zeraix/zeraix

Open-source local AI workspace — advancing on-device inference.

repo
NightMean/OlliteRT

Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source

repo
Mobile-Artificial-Intelligence/llama_sdk

lcpp is a dart implementation of llama.cpp used by the mobile artificial intelligence distribution (maid)

repo
alibaba/InferSim

A Lightweight LLM Inference Performance Simulator

repo
thomas9120/LLama-GUI

User friendly GUI for configuring and launching llama.cpp

repo
Tylogi/TyloQuant

Get more intelligence from every bit. Better quantization formats and smarter calibration let larger, stronger models run smoothly on the hardware you already o…

repo
IlyaGusev/saiga_bot

Telegram bot for different language models. Supports system prompts and images

repo
john-rocky/apple-silicon-llm-bench

Neutral, reproducible benchmark for local LLMs on Apple Silicon (Mac · iPhone · iPad) — MLX, llama.cpp, CoreML, Apple Foundation Models

repo
MADEVAL/FerryAI

Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process via FFI. Chat, streaming, embeddings, RAG and vect…

repo
CharlesPikachu/FreeGPTHub

FreeGPTHub: A truly free, unified GPT API gateway—unlike “fake-free” alternatives that hide paywalls behind quotas, trials, or mandatory top-ups. (真正免费的GPT统一接口,…

repo
Xyntopia/taskyon

Browser based Interface for Generative AI. Chat/Agent/Taskmanager Hybrid.

Related topics