jeinlee1991/chinese-llm-benchmark
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan →
“Fuzzy title match (0.73): “Live Gurbani Tracking: A Benchmark and Reference System for ” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 59%HoloCount: A Holistic Visual Counting Benchmark for MLLMs →
“Fuzzy title match (0.73): “HoloCount: A Holistic Visual Counting Benchmark for MLLMs” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 59%EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage →
“Fuzzy title match (0.73): “EgoPolice: A Benchmark for Egocentric Video Understanding in” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 59%SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models →
“Fuzzy title match (0.73): “SPEARBench: A Benchmark for Naturalness Evaluation in Stream” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 87%MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation →
“Fuzzy title match (0.94): “MedRealMM: A Real-World Multimodal Benchmark for Chinese Onl” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 59%EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots →
“Fuzzy title match (0.73): “EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Ev” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 59%Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering →
“Fuzzy title match (0.73): “Clinically Structured Rank-Gated LoRA for Cross-Benchmark Me” ≈ “jeinlee1991/chinese-llm-benchmark””
- FuzzySimilar title/name (fuzzy) · 59%The Human Creativity Benchmark →
“Fuzzy title match (0.73): “The Human Creativity Benchmark” ≈ “jeinlee1991/chinese-llm-benchmark””
