Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · yesterday

Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning

Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical reasoning on the Countdown task. We first analyze reasoning failures and find that calculation errors account for a substantial portion of incorrect responses. We then construct supervised fine-tuning datasets to teach the model useful tool-use patterns and how to interpret returned outputs. Building on this tool-formatted policy, we apply several on-policy r

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • Fuzzy title match (0.76): “Learning to Use Tools: Reinforcement Learning for Tool-Integ” ≈ “MathFoundationRL/Book-Mathematical-Foundation-of-Reinforceme”

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Minghui Xu

    Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning

  • LinkedLinked via arxiv author · 85%Jiazi Wang

    Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning

  • FuzzySimilar title/name (fuzzy) · 84%amitness/learning

    Fuzzy title match (0.92): “Learning to Use Tools: Reinforcement Learning for Tool-Integ” ≈ “amitness/learning”

Implements (incoming)

authored (incoming)

Related across the graph

Topics