Topic

Llama Cpp

46 items across the graph — tagged with Llama Cpp.

From the graph · 46

repo
SciSharp/LLamaSharp

A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.

repo
Osmantic/ODS

Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.

repo
AtomicBot-ai/atomic-agent

Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.

repo
antoinezambelli/forge

A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows

repo
mybigday/llama.rn

React Native binding of llama.cpp

repo
kennss/SiliconScope

Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking

repo
ciddwd/overlay-translator

无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, vis…

repo
hogeheer499-commits/strix-halo-guide

Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.

repo
raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

repo
OleksandrChekhovskyi/hax

A minimalist, terminal-native coding agent written in C.

repo
mohitsoni48/TurboLLM

Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No E…

repo
Scottcjn/ram-coffers

NUMA-distributed weight banking for LLM inference on IBM POWER8. 147 t/s (8.8x stock). Part of the Proof of Physical AI stack.

repo
GaeaRuiW/kube-llmops
repo
engeldlgado/toshllm

Run large language models locally on Intel Macs with AMD GPUs - native macOS app with Metal acceleration

repo
raiyanyahya/llmaker

Selfhost modern LLM stacks. Run the whole fleet from your terminal

repo
swellweb/reame

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a…

repo
FedericoTs/quantprobe

Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hard…

repo
thomas9120/LLama-GUI

User friendly GUI for configuring and launching llama.cpp

repo
Helldez/BigMoeOnEdge

Run a 120B-parameter MoE (60 GB) on a 12 GB phone. CPU-only, lossless, on stock llama.cpp

repo
aaronnat23/disp8ch

Self-hosted AI workspace where chat becomes visual workflows, multi-agent operations, and reviewable automations. Local memory; local or cloud models

repo
Hal0ai/hal0

Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI + systemd.

repo
fengzhizi715/OpenVitamin

OpenVitamin is a local-first AI execution platform that unifies Agents, Workflows, and multi-model inference into a single programmable system — designed for bu…

repo
Scottcjn/llama-cpp-tigerleopard

WORLD FIRST: llama.cpp for Mac OS X Tiger & Leopard on PowerPC G4/G5

repo
ArkaneFans/servllama

1-Click LLM Server on Your Phone — no Termux needed! 无需Termux,一键让你的手机变成LLM服务器!

repo
Aatricks/llmedge

Android native AI inference library, bringing text, image, video, STT, TTS inference

repo
john-rocky/apple-silicon-llm-bench

Neutral, reproducible benchmark for local LLMs on Apple Silicon (Mac · iPhone · iPad) — MLX, llama.cpp, CoreML, Apple Foundation Models

repo
MADEVAL/FerryAI

Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process via FFI. Chat, streaming, embeddings, RAG and vect…

repo
hadihonarvar/flock

Self-hosted LLM gateway. One Go binary turns your Macs and Linux boxes into a private inference cluster — multi-machine routing, sharding via llama.cpp-RPC, per…

repo
off-grid-ai/OGAD

Off Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API k…

repo
Etamus/NeveAI

Neve AI é uma plataforma de IA local privacy-first, desenvolvida para oferecer uma experiência de alta performance na execução de LLMs, reduzindo a dependência…

repo
modelship-ai/modelship

Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and im…

repo
alez007/modelship

Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and im…

repo
DaoyuanLi2816/llm-gpu-lab

One GPU. Full LLM workflow. Real benchmarks. No cloud required.

repo
yatesdr/go-llm-proxy

Lightweight proxy for LLM

repo
kekzl/imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reason…

repo
mnemosyne-systems/orangu

Advanced code editor using local AI

repo
yumiaura/aicommit

AI commit messages from your local LLM - Ollama / llama-cpp backed, offline, Conventional Commits, with review + changelog modes.

repo
notwitcheer/llm-bench-rig

Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards.

repo
off-grid-ai/off-grid-ai-desktop

Off Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API k…

repo
gdevenyi/huggingface-estimate

A web-based memory usage and performance calculator for Huggingface GGUF models

repo
Fangyuan025/Chaty

Private, on-device AI desktop app — GGUF (llama.cpp) & MLX models, a local coding agent, RAG knowledge base, Deep Research, vision and voice. 100% offline, no a…

repo
allanschramm/local-model-autotuning

Este repositório visa ensinar o usuário a como usar IAs localmente da forma correta e encontrar a melhor configuração pro modelo escolhido de forma automática e…

repo
matthewdcage/llm-swarm-router

Run the LLM Swarm Router on machines to distribute Local Ai to the Swarm - More Machines - MORE SPEED

repo
aivrar/multi-turboquant

Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cach…

repo
fewtarius/llama-ai

A project repository for work on improving local LLMs on my personal AMD devices

repo
luojieLLMaaS/haxiv

Deep learning framework for LLMs (Llama/Gemma/Qwen) on CPU, Apple MLX, Metal, CUDA. Load PyTorch/ONNX/TF/GGUF with zero conversion. PyTorch alternative with nat…

Related topics