repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
patrick-toulme/harnessgym
Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay fresh sessions with it activated. Works with Codex and Claude Code.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%AgentCore-8B →
- PossiblePossibly related (embedding) · 52%I built an agent Harness for Small Models. I got Qwen 3.5 4b managing servers. →
- PossiblePossibly related (embedding) · 50%Agentic Hardware Design as Repository-Level Code Evolution →
- PossiblePossibly related (embedding) · 50%AutoTrainess: Teaching Language Models to Improve Language Models Autonomously →
- PossiblePossibly related (embedding) · 50%Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents →
- PossiblePossibly related (embedding) · 56%Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study →
- PossiblePossibly related (embedding) · 53%CurateEvo: Data-Curation Evolving for Agentic Post-Training →
- PossiblePossibly related (embedding) · 65%How self-improving harnesses are rewriting the agent engineering playbook - TechTalks →
Related to
Covers
Implements
Implements (incoming)
Covers (incoming)
Related across the graph
newsTraining a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]newsA primer on self-improving agent harnesses - SubstackpaperCurateEvo: Data-Curation Evolving for Agentic Post-TrainingnewsI built an agent Harness for Small Models. I got Qwen 3.5 4b managing servers.newsHow self-improving harnesses are rewriting the agent engineering playbook - TechTalkspaperLearning from Failure: Inference-Time Self-Improvement for Computer-Use AgentspaperReasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational studymodelAgentCore-8BpaperAutoTrainess: Teaching Language Models to Improve Language Models AutonomouslypaperAgentic Hardware Design as Repository-Level Code Evolution
