Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exploratory insight discovery. We introduce DataGovBench, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios. The benchmark includes two tasks: Table QA that requires solving complex decomposable questions and producing textual answers or visualizations, and Table Insight that
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 52%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 52%douglasjordan2/c0 →
- PossiblePossibly related (embedding) · 49%Tencent/WeKnora →
- PossiblePossibly related (embedding) · 49%spiceai/spiceai →
- LinkedLinked via arxiv author · 85%So Hasegawa →
“Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities”
- LinkedLinked via arxiv author · 85%Shailaja Keyur Sampat →
“Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities”
- LinkedLinked via arxiv author · 85%Lei Liu →
“Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities”
