BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contain
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%jeinlee1991/chinese-llm-benchmark →
“Fuzzy title match (0.73): “BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Be” ≈ “jeinlee1991/chinese-llm-benchmark””
- LinkedLinked via arxiv author · 85%Liubov Chubarova →
“BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models”
- LinkedLinked via arxiv author · 85%Alexandra Kuleshova →
“BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models”
- LinkedLinked via arxiv author · 85%Daniil Volkov →
“BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models”
- LinkedLinked via arxiv author · 85%Kirill Sultanov →
“BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models”
- LinkedLinked via arxiv author · 85%Alexey Zaytsev →
“BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models”
- PossiblePossibly related (embedding) · 51%Enterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU - blog.google →
