Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 13h ago

jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Implements

paperLive Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh KirtanpaperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperEgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera FootagepaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperEMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support ChatbotspaperClinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question AnsweringpaperThe Human Creativity BenchmarkpaperBeyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkpaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionpaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperAbsoluteDegradation: A Physics-Inspired Synthetic Film-Degradation Pipeline and Archival Film Restoration BenchmarkpaperAdversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy AmbiguitypaperMM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationpaperMedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionpaperDoes generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific LiteraturepaperFootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information GamespaperFrontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoningpaperVecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector FontspaperEvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingpaperVEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester DesignpaperSpEmoC: A Balanced Speaker-Segment Multimodal Emotion BenchmarkpaperMIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical ExamspaperExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesispaperTwo-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenesspaperHalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question AnsweringpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperRUMBA: Russian User Memory BenchmarkpaperOne More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification PoliciespaperFuture Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed WindowpaperEdit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZpaperAnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMspaperReconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication BibliographiespaperKrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisorypaperEduArt: An educational-level benchmark for evaluating art history knowledge in large language modelspaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperKnowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMspaperArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text RenderingpaperGenerating Benchmark Health Data Using a Tabular Diffusion TransformerpaperGBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection

Covers (incoming)

Related across the graph

paperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperEgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera FootagepaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperEMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support ChatbotspaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperClinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question AnsweringpaperKnowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMspaperGenerating Benchmark Health Data Using a Tabular Diffusion TransformerpaperThe Human Creativity BenchmarkpaperBeyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkpaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionpaperDoes generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific LiteraturepaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperAbsoluteDegradation: A Physics-Inspired Synthetic Film-Degradation Pipeline and Archival Film Restoration BenchmarkpaperMM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationpaperGBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack DetectionpaperMedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperKrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisorypaperRUMBA: Russian User Memory BenchmarkpaperAdversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy AmbiguitypaperOne More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification PoliciesnewsPrism-ML's Bonsai-27B BenchmarkspaperAnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMspaperExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesispaperArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text RenderingpaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperEvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingpaperEduArt: An educational-level benchmark for evaluating art history knowledge in large language modelspaperHalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question AnsweringpaperTwo-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenesspaperEdit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZpaperLive Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh KirtanpaperReconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication BibliographiespaperFrontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoningpaperVecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector FontspaperMIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical ExamspaperFuture Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed WindowpaperVEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester DesignpaperSpEmoC: A Balanced Speaker-Segment Multimodal Emotion BenchmarkpaperFootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

Topics