Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /jeinlee1991/chinese-llm-benchmark
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 28d ago

jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan →

    “Fuzzy title match (0.73): “Live Gurbani Tracking: A Benchmark and Reference System for ” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 59%HoloCount: A Holistic Visual Counting Benchmark for MLLMs →

    “Fuzzy title match (0.73): “HoloCount: A Holistic Visual Counting Benchmark for MLLMs” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 59%EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage →

    “Fuzzy title match (0.73): “EgoPolice: A Benchmark for Egocentric Video Understanding in” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 59%SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models →

    “Fuzzy title match (0.73): “SPEARBench: A Benchmark for Naturalness Evaluation in Stream” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 87%MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation →

    “Fuzzy title match (0.94): “MedRealMM: A Real-World Multimodal Benchmark for Chinese Onl” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 59%EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots →

    “Fuzzy title match (0.73): “EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Ev” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 59%Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering →

    “Fuzzy title match (0.73): “Clinically Structured Rank-Gated LoRA for Cross-Benchmark Me” ≈ “jeinlee1991/chinese-llm-benchmark””

  • FuzzySimilar title/name (fuzzy) · 59%The Human Creativity Benchmark →

    “Fuzzy title match (0.73): “The Human Creativity Benchmark” ≈ “jeinlee1991/chinese-llm-benchmark””

Implements

paperLive Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh KirtanpaperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperEgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera FootagepaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperEMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support ChatbotspaperClinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question AnsweringpaperThe Human Creativity BenchmarkpaperBeyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkpaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionpaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperAbsoluteDegradation: A Physics-Inspired Synthetic Film-Degradation Pipeline and Archival Film Restoration BenchmarkpaperAdversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy AmbiguitypaperMM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationpaperMedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionpaperDoes generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific LiteraturepaperKrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisorypaperEduArt: An educational-level benchmark for evaluating art history knowledge in large language modelspaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperKnowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMspaperArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text RenderingpaperFootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information GamespaperFrontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoningpaperVecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector FontspaperEvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingpaperVEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester DesignpaperSpEmoC: A Balanced Speaker-Segment Multimodal Emotion BenchmarkpaperMIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical ExamspaperExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesispaperTwo-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenesspaperHalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question AnsweringpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperRUMBA: Russian User Memory BenchmarkpaperOne More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification PoliciespaperFuture Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed WindowpaperEdit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

Covers (incoming)

newsPrism-ML's Bonsai-27B Benchmarks

Related across the graph

paperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperEgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera FootagepaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperEMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support ChatbotspaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperClinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question AnsweringpaperKnowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMspaperThe Human Creativity BenchmarkpaperBeyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkpaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionpaperDoes generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific LiteraturepaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperAbsoluteDegradation: A Physics-Inspired Synthetic Film-Degradation Pipeline and Archival Film Restoration BenchmarkpaperMM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationpaperMedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperKrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisorypaperRUMBA: Russian User Memory BenchmarkpaperAdversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy AmbiguitypaperOne More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification PoliciesnewsPrism-ML's Bonsai-27B BenchmarkspaperExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesispaperArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text RenderingpaperCross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation TrainingpaperEvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingpaperEduArt: An educational-level benchmark for evaluating art history knowledge in large language modelspaperHalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question AnsweringpaperTwo-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenesspaperEdit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZpaperLive Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh KirtanpaperFrontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoningpaperVecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector FontspaperMIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical ExamspaperFuture Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed WindowpaperVEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester DesignpaperSpEmoC: A Balanced Speaker-Segment Multimodal Emotion BenchmarkpaperFootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
Knowledge path·PHoloCount: A Holistic Visual Counting Benchmark for MLLMs→PEgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage→PSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models→Rjeinlee1991/chinese-llm-benchmark

Topics

agentic-aiartificial-intelligencellm-agentllm-evaluation

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Graph trust82Primary
Graph score6300