StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies,
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Jinghan Tan →
“StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models”
- LinkedLinked via arxiv author · 85%Yuanzheng Wang →
“StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models”
- LinkedLinked via arxiv author · 85%Lu Chen →
“StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models”
- LinkedLinked via arxiv author · 85%Zijun Chen →
“StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models”
- LinkedLinked via arxiv author · 85%Yuqian Wang →
“StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models”
