Read original ↗
paperarXivTrust 82 · PrimaryPublished yesterdayLive · 1h ago

MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and revise reward functions as monolithic programs, making it difficult to reliably preserve and reuse effective components discovered in earlier iterations, leading to unstable performance across iterations. To address this, we propose Module Level Reward Evolution Framework (MLREF). At the core of MLREF is a module pool, a persistent repository of reusable reward components. MLREF treats the module pool as the primary op

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning

    Fuzzy title match (0.73): “MLREF: Efficient Module Reuse for Reward Design in Reinforce” ≈ “aymericdamien/TopDeepLearning”

  • LinkedLinked via arxiv author · 85%Chenglin Liu

    MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

  • LinkedLinked via arxiv author · 85%Weixun Wang

    MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

  • LinkedLinked via arxiv author · 85%Ruishuo Chen

    MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

  • LinkedLinked via arxiv author · 85%Zhuoran Li

    MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

Implements (incoming)

authored (incoming)

Related across the graph

Topics