repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 6d ago
AgileRL/AgileRL
Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimization.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%RL without TD learning →
- PossiblePossibly related (embedding) · 58%Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training →
- PossiblePossibly related (embedding) · 57%Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline) →
- PossiblePossibly related (embedding) · 53%Joint Learning of Experiential Rules and Policies for Large Language Model Agents →
- PossiblePossibly related (embedding) · 52%Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models →
- PossiblePossibly related (embedding) · 47%Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design →
- PossiblePossibly related (embedding) · 57%SCOPE-RL: Optimizing Reasoning Paths Before and After Success →
- PossiblePossibly related (embedding) · 55%Active Offline-to-Online Reinforcement Learning →
Covers
Implements
paperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action Models
Implements (incoming)
paperImproving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function DesignpaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperActive Offline-to-Online Reinforcement LearningpaperTransformer-Guided Swarm Intelligence for Frugal Neural Architecture SearchpaperA Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous ManipulationpaperGradient-free learning of a closed-loop wall controller for turbulent drag reductionpaperLearning-enabled Acceleration of Scenario-based Model Predictive ControlpaperVerifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage ControlpaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperKnowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision ProcessespaperSILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable RoutingpaperAdaptive Inference Batching using Policy GradientspaperTREK: Distill to Explore, Reinforce to RefinepaperWeak-to-Strong Generalization via Direct On-Policy Distillation
Related across the graph
paperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperVerifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage ControlpaperImproving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function DesignnewsRL without TD learningpaperAdaptive Inference Batching using Policy GradientspaperActive Offline-to-Online Reinforcement LearningpaperSILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable RoutingpaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperWeak-to-Strong Generalization via Direct On-Policy DistillationpaperA Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its MechanismpaperKnowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision ProcessespaperLearning-enabled Acceleration of Scenario-based Model Predictive ControlpaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperTransformer-Guided Swarm Intelligence for Frugal Neural Architecture SearchpaperA Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous ManipulationpaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperGradient-free learning of a closed-loop wall controller for turbulent drag reductionpaperTREK: Distill to Explore, Reinforce to Refine
