Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. We ask whether a unified LMM can improve both abilities autonomously using only unlabeled images. We propose a self-evolving training framework with three internal roles: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that synthesizes images. Training uses only self-derived consistency signals, without human annotatio
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownVioletVision-3B →
- LinkedLinked via unknownHow Preply combines AI and human tutors to personalize learning →
- LinkedLinked via unknownvlm-starter →
- PossiblePossibly related (embedding) · 64%EvolvingLMMs-Lab/LLaVA-OneVision-2 →
- PossiblePossibly related (embedding) · 49%redai-infra/Relax →
- PossiblePossibly related (embedding) · 46%LDJ-creat/video-helper →
- PossiblePossibly related (embedding) · 48%lightly-ai/lightly →
