repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
swellweb/reame
A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It →
- PossiblePossibly related (embedding) · 51%Any good uses for a 192 GB DDR3 Server in the LLM world? →
- PossiblePossibly related (embedding) · 48%What machine is best for my setup? [D] →
- PossiblePossibly related (embedding) · 59%7 Best Self-Hosted Inference Servers for Open-Source Models, Compared (2026) - HackerNoon →
- PossiblePossibly related (embedding) · 47%Why goodput matters more than throughput for LLM serving →
- PossiblePossibly related (embedding) · 50%Torrents arrived →
Covers
Covers (incoming)
Related across the graph
newsWhat machine is best for my setup? [D]newsAny good uses for a 192 GB DDR3 Server in the LLM world?newsTorrents arrivednewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix Itnews7 Best Self-Hosted Inference Servers for Open-Source Models, Compared (2026) - HackerNoonnewsWhy goodput matters more than throughput for LLM serving
