SPyCE: Skill-Policy Co-evolution for Multimodal Agents
Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps. Existing reinforcement learning methods reduce trajectories to scalar rewards, forcing the policy to discover reusable tool-use patterns from scratch on every new task; memory-based alternatives retain past experience, yet they rely on test-time retrieval, without updating the policy to absorb reusable patterns from that experience. Our key insight is that multimodal reasoning trajectories should be distilled into reusable skills that co-evolve with the policy during training, ra
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Ru Zhang →
“SPyCE: Skill-Policy Co-evolution for Multimodal Agents”
- LinkedLinked via arxiv author · 85%Weijie Qiu →
“SPyCE: Skill-Policy Co-evolution for Multimodal Agents”
- PossiblePossibly related (embedding) · 55%Exploring continual learning without replay buffers: Our findings using dynamic task-similarity routing [P] →
