repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
Prism-Shadow/GDPevo
A Benchmark for Evaluating Agent Self-Evolution on Real Business Tasks
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution →
- PossiblePossibly related (embedding) · 57%Self-Evolving World Models for LLM Agent Planning →
- PossiblePossibly related (embedding) · 54%NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations →
- PossiblePossibly related (embedding) · 52%EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments →
- PossiblePossibly related (embedding) · 52%TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution →
- PossiblePossibly related (embedding) · 56%The real cost, security, and culture problems behind enterprise AI agents →
- PossiblePossibly related (embedding) · 58%The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway →
- PossiblePossibly related (embedding) · 54%Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them →
Implements
paperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperSelf-Evolving World Models for LLM Agent PlanningpaperEvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive EnvironmentspaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
Covers
Covers (incoming)
Implements (incoming)
Related across the graph
paperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperSelf-Evolving World Models for LLM Agent PlanningpaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentsnewsThe real cost, security, and culture problems behind enterprise AI agentsnewsEnterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify themnewsNVIDIA Brings Trusted, 24/7 AI Agents to Telecom OperationspaperEvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive EnvironmentspaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionnewsThe agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
