newsReddit r/LocalLLaMATrust 52 · CommunityPublished 1mo agoLive · 1mo ago
I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads
I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads — prefill dominates everything, and KV head count beats parameter count I've been running local LLMs for agentic workflows (tool use, coding agents, RAG) and kept seeing people obsess over tg128 (token generation speed) as the headline performance metric. So I ran a structured long-context benchmark to figure out what actually matters when your c
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 56%TraceLab: Characterizing Coding Agent Workloads for LLM Serving →
- PossiblePossibly related (embedding) · 54%Evaluate a model properly →
- PossiblePossibly related (embedding) · 51%Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
- PossiblePossibly related (embedding) · 51%Context-Engine-AI/Context-Engine →
- PossiblePossibly related (embedding) · 51%harvard-cns/orla →
- PossiblePossibly related (embedding) · 50%agentjido/llm_db →
- PossiblePossibly related (embedding) · 54%LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget →
Covers
Covers (incoming)
Related across the graph
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperTraceLab: Characterizing Coding Agent Workloads for LLM ServingrepoTura-AI/turarepogqgs/llm100kbenchrepoagentjido/llmdbrepoContext-Engine-AI/Context-Enginerepoharvard-cns/orlapaperPACE: A Proxy for Agentic Capability Evaluationrepoagentjido/llm_dbtutorialEvaluate a model properlypaperLongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU BudgetpaperTeach it to stop, not just to clickpaperWindowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context
