Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 3d agoLive · 3d ago

Has anyone actually made 64k feel like 300k+ with recursive local agents?

I'm running Qwen 3.8 27B locally on a single GPU. I can push the context to 131k, but I'd rather run it faster at 64k if the agent can manage context properly. What I have in mind is pretty simple: one model stays loaded the whole time main agent gets 64k when something is too big, it spawns a fresh child with only the task and context it needs if that child gets a 100k document, it can split the job again or spawn

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Related across the graph