repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
teilomillet/retrain
a Python library that uses Reinforcement Learning (RL) to train LLMs.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training →
- PossiblePossibly related (embedding) · 50%Reinforcement Learning without Ground-Truth Solutions can Improve LLMs →
- PossiblePossibly related (embedding) · 49%Would having a dedicated programming language specifically for LLMs be a viable solution? [D] →
- PossiblePossibly related (embedding) · 48%I shrank a transformer until every number fitted on the screen and made the weights editable [R] →
- PossiblePossibly related (embedding) · 47%Generative Skill Composition for LLM Agents →
- PossiblePossibly related (embedding) · 47%Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design →
Implements
Covers
Implements (incoming)
Related across the graph
paperImproving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function DesignnewsWould having a dedicated programming language specifically for LLMs be a viable solution? [D]paperGenerative Skill Composition for LLM AgentsnewsI shrank a transformer until every number fitted on the screen and made the weights editable [R]paperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperReinforcement Learning without Ground-Truth Solutions can Improve LLMs
