Topic

Gguf

50 items across the graph — tagged with Gguf.

From the graph · 50

repo
AlexsJones/llmfit

Hundreds of models & providers. One command to find what runs on your hardware.

repo
Michael-A-Kuykendall/shimmy

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

repo
Andyyyy64/whichllm

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it ins…

model
HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

Hugging Face model with 3444 likes. Tags: gguf, uncensored, qwen3.6, moe, vision, multimodal, image-text-to-text, en, zh, multilingual

model
google/gemma-7b

Hugging Face model with 3397 likes. Tags: transformers, safetensors, gguf, gemma, text-generation, arxiv:2305.14314, arxiv:2312.11805, arxiv:2009.03300, arxiv:1…

repo
withcatai/node-llama-cpp

Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level

repo
sammcj/gollama

Go manage your Ollama models

repo
MakazhanAlpamys/Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

repo
kitops-ml/kitops

An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

repo
AtomicBot-ai/Atomic-Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.

repo
eastriverlee/LLM.swift

LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.

repo
AtomicBot-ai/atomic-agent

Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.

repo
jegly/Box

The most advanced, fully offline client-side AI suite on Android today.

repo
ddalcu/mlx-serve

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling…

repo
hybridgroup/yzma

Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.

repo
lone-cloud/gerbil

A desktop app for running Large Language Models locally.

repo
ciddwd/overlay-translator

无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, vis…

repo
avifenesh/memra

memra is a Rust + CUDA inference engine built for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. It serves GGUF models over an OpenAI-compatible API, and every def…

repo
hogeheer499-commits/strix-halo-guide

Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.

repo
beehive-lab/GPULlama3.java

GPU-accelerated Llama3.java inference in pure Java using TornadoVM.

repo
raketenkater/ggrun

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuni…

repo
mgonzs13/llama_ros

llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2

repo
zhongkaifu/TensorSharp

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP…

repo
jjang-ai/jangq

JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon

repo
giannisanni/pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-…

repo
defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM,…

repo
mohitsoni48/TurboLLM

Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No E…

repo
Mobile-Artificial-Intelligence/llama_sdk

lcpp is a dart implementation of llama.cpp used by the mobile artificial intelligence distribution (maid)

repo
timtoole02/Camelid

Camelid: a Rust-native local inference backend with evidence-gated model compatibility.

repo
swellweb/reame

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a…

repo
docusealco/rllama

Ruby FFI bindings for llama.cpp to run open-source LLMs such as GPT-OSS, Qwen 3.5, Gemma 4, and Llama 3 locally with Ruby.

repo
noumena-labs/Sipp

AI inference, packed simply. A blazing-fast, zero-dependency WebGPU runtime to run GGUF models directly in the browser. Features a symmetric API for seamless lo…

repo
llamastash/llamastash

A fast terminal native app (TUI) and CLI with init wizard for launching local LLMs with zero overhead

repo
FedericoTs/quantprobe

Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hard…

repo
AarambhDevHub/aarambh-studio

🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/docu…

repo
Helldez/BigMoeOnEdge

Run a 120B-parameter MoE (60 GB) on a 12 GB phone. CPU-only, lossless, on stock llama.cpp

repo
Entrpi/ds4-on-spark

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~2x upstream prefill, ~1.5x decode, DSpark, and full continuous…

repo
Aatricks/llmedge

Android native AI inference library, bringing text, image, video, STT, TTS inference

repo
linux4life1/front-porch-AI

Local-first AI character chat & roleplay for Windows/macOS/Linux — a living Realism Engine, built-in TTS & image generation, and The Stoop community character h…

repo
hadihonarvar/flock

Self-hosted LLM gateway. One Go binary turns your Macs and Linux boxes into a private inference cluster — multi-machine routing, sharding via llama.cpp-RPC, per…

repo
MADEVAL/FerryAI

Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process via FFI. Chat, streaming, embeddings, RAG and vect…

repo
eullm/eullm

Open-source platform for creating, distributing and running sovereign EU-compliant LLMs. Verticalize any model for your domain, language and brand. AI Act ready…

repo
off-grid-ai/OGAD

Off Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API k…

repo
thomas9120/LLama-GUI

User friendly GUI for Llama.cpp for easy configuration and launching.

repo
DaoyuanLi2816/llm-gpu-lab

One GPU. Full LLM workflow. Real benchmarks. No cloud required.

repo
Sudharsanselvaraj/Token-Print

Interactive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.

repo
kekzl/imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reason…

repo
AarambhDevHub/aarambh-ai

🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/docu…

repo
Yog-Sotho/LLM-fine-tuner

Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF exp…

repo
off-grid-ai/off-grid-ai-desktop

Off Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API k…

Related topics