repoGitHubTrust 82 · PrimaryPublished 11d agoLive · 11d ago
zhangxjohn/LLM-Agent-Benchmark-List
A banchmark list for evaluation of large language models.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- PossiblePossibly related (embedding) · 59%Large Language Models: Qwen3 Offers AI Models For Deeper Reasoning And Faster Responses - Trend Hunter →
