repoGitHubTrust 82 Β· PrimaryPublished 1mo agoLive Β· 12d ago
ManasVardhan/bench-my-llm
ποΈ Dead-simple LLM benchmarking CLI - latency, cost, and quality metrics
Lineage graph
Paper β model β repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it β so bad links are debuggable.
- PossiblePossibly related (embedding) Β· 63%Evaluate a model properly β
- PossiblePossibly related (embedding) Β· 61%Why goodput matters more than throughput for LLM serving β
- PossiblePossibly related (embedding) Β· 55%Are there good closed vs open LLM rankings? Also, are 70Bβ350B models actually worth it? β
- PossiblePossibly related (embedding) Β· 55%Atomicwork And New Measure Open Source AI Benchmark For ITSM - Open Source For You β
- PossiblePossibly related (embedding) Β· 51%Docker Performance Benchmarking 2026: Optimization Standards Gap - technosports.co.in β
- PossiblePossibly related (embedding) Β· 58%Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together - MarkTechPost β
- PossiblePossibly related (embedding) Β· 51%Docker Performance Benchmarking 2026: Optimization Standards Gap - TechnoSports Media Group β
- PossiblePossibly related (embedding) Β· 51%ChatGPT Claude Docker Benchmarks: Performance Analysis - TechnoSports Media Group β
Related to
Covers
Covers (incoming)
newsAtomicwork And New Measure Open Source AI Benchmark For ITSM - Open Source For YounewsDocker Performance Benchmarking 2026: Optimization Standards Gap - technosports.co.innewsLiquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together - MarkTechPostnewsDocker Performance Benchmarking 2026: Optimization Standards Gap - TechnoSports Media GroupnewsChatGPT Claude Docker Benchmarks: Performance Analysis - TechnoSports Media GroupnewsBenchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6 - Amazon Web Services (AWS)newsMeasuring LLM performance drift: observations and methodology from 31,352 repeated benchmark measurements [D]
Related across the graph
newsMeasuring LLM performance drift: observations and methodology from 31,352 repeated benchmark measurements [D]newsLiquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together - MarkTechPostnewsChatGPT Claude Docker Benchmarks: Performance Analysis - TechnoSports Media GroupnewsDocker Performance Benchmarking 2026: Optimization Standards Gap - technosports.co.innewsAre there good closed vs open LLM rankings? Also, are 70Bβ350B models actually worth it?newsBenchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6 - Amazon Web Services (AWS)tutorialEvaluate a model properlynewsWhy goodput matters more than throughput for LLM servingnewsAtomicwork And New Measure Open Source AI Benchmark For ITSM - Open Source For YounewsDocker Performance Benchmarking 2026: Optimization Standards Gap - TechnoSports Media Group
