newsReddit r/LocalLLaMATrust 52 · CommunityPublished 24d agoLive · 24d ago
CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful
I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.cpp focused on a problem that matters a lot on slower hardware: repeated prompt processing. Not only does it have a new "SSD" based cache, but it also has some other improvements with cach
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%cnighswonger/claude-code-cache-fix →
- PossiblePossibly related (embedding) · 53%vcache-project/vCache →
- PossiblePossibly related (embedding) · 50%test5630352/llm-cache-optimize →
- PossiblePossibly related (embedding) · 50%pythongiant/KVBoost →
- PossiblePossibly related (embedding) · 49%xcena-dev/maru →
- PossiblePossibly related (embedding) · 46%EdgarOrtegaRamirez/llm-cache →
