Local Llm
47 items across the graph — tagged with Local Llm.
From the graph · 47
Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no clou…
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separatio…
A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
Build AI agents from first principles using a local LLM - no frameworks, no cloud APIs, no hidden reasoning.
🚀 Self-hosted AI automation platform. Deploy n8n, Ollama, Flowise, RAG, Supabase & 30+ tools with one command. Auto HTTPS. Free Zapier/Make alternative.
Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking
Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.
The future of AI is local. Time to build your own Personal AI Computer.
Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
AI-native IDE with agentic workflows, iPhone emulation on Windows/Linux, PyTorch ML Studio, and ROCm-optimized local AI. Built for security researchers and cros…
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No E…
A minimalist, terminal-native coding agent written in C.
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-…
Self-hostable multimodal chat with local LLMs (Ollama/OpenAI): PDF RAG, image chat, and Whisper voice, Streamlit + Docker.
Orchestrate a swarm of Claude Code agents with a local brain that learns from you.
M-Courtyard: Local AI Model Fine-tuning Assistant for Apple Silicon. Zero-code, zero-cloud, privacy-first desktop app powered by Tauri + React + mlx-lm.
Run Open Source/Open Weight LLMs locally with OpenAI compatible APIs
Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source
Synthetic Autonomic Mind - An AI assistant for everyone.
Your AI intranet: network the computers you already own for inference and training.
A simpler self-hosted alternative to Open WebUI. Bring your own API keys or local models. Native Android client in closed beta.
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process via FFI. Chat, streaming, embeddings, RAG and vect…
Frontier models sell confidence. FABULA ships proof — an agent harness where any model is a swappable chip and every finished run mints a replayable, context-fi…
A memory-first AI agent that remembers why decisions were made — not just the last message. Runs local (Ollama), cloud (Claude · OpenAI · Gemini), or decentrali…
An AI-native OS. Models run on your own hardware, nodes federate peer-to-peer with no broker, and the API is OpenAI-compatible. Boots, has its own Wayland compo…
Open-source platform for creating, distributing and running sovereign EU-compliant LLMs. Verticalize any model for your domain, language and brand. AI Act ready…
Off Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API k…
Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.
Cortex is a fast, private desktop AI assistant for running local Large Language Models with Ollama. Everything stays on your device: no cloud, no external serve…
An easy-to-use, fast toolkit to scale up RL post-training on a single node.
LLM infradebugging and diagnostic tool
Run LLMs, VLMs, ASR, TTS, diarization and more fully on-device with Apple's Core AI framework (iOS/macOS 27) — one line of Swift per model, 53 models pinned to…
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
🤵 AIfred-Intelligence — self-hosted Multi-Agent Assistant with Debate Modes (Symposion/Tribunal), Voice (STT + Streaming-TTS), RAG with Long-Term Memory, Web R…
An LLM-powered social simulation engine
Chain two LLMs — local or cloud: a Prompter that refines your idea, and a Coder that writes the code.
Off Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API k…
Swift 6 agent SDK: type-safe tools, streaming, cloud + on-device inference via MLX on Apple Silicon
Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GP…
Make local LLM inference faster with chunk-level KV cache reuse
Run the LLM Swarm Router on machines to distribute Local Ai to the Swarm - More Machines - MORE SPEED
One Mac process. Many models. Real speed. Multi-model LLM serving with prefix reuse, MTP acceleration, and OpenAI APIs — built for Apple Silicon, measured again…
Este repositório visa ensinar o usuário a como usar IAs localmente da forma correta e encontrar a melhor configuração pro modelo escolhido de forma automática e…
End-to-end agentic quantum-ML pipeline: PennyLane GPU quantum-kernel anomaly detection + LangGraph multi-agent system + locally-hosted Qwen3-14B, benchmarked ac…
