Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 5d ago

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness p

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B

    Fuzzy title match (0.73): “Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Lea” ≈ “AgentCore-8B”

  • FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent

    Fuzzy title match (0.94): “Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Lea” ≈ “SWE-agent/SWE-agent”

  • FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent

    Fuzzy title match (0.94): “Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Lea” ≈ “zhayujie/CowAgent”

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzySimilar title/name (fuzzy) · 59%NousResearch/hermes-agent

    Fuzzy title match (0.73): “Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Lea” ≈ “NousResearch/hermes-agent”

  • LinkedLinked via arxiv author · 85%Zixuan Wang

    Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

  • LinkedLinked via arxiv author · 85%Yanrui Miao

    Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics