VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following
Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feedback from failed samples, verifier outcomes, and target-model errors. We present VISA (Visual Instruction Synthesis Agent), an agentic framework that reformulates multimodal instruction synthesis as a self-evolving loop. At each round, VISA analyzes an image to filter incompatible constraints and discover new verifiable ones, samples diversity- and difficulty-aware constraint s
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Min Zeng →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Guanxin Tan →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Libin Cen →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Yawei Wen →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Chuanrui Hu →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Liuyang Bian →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Xiaolong Chen →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
- LinkedLinked via arxiv author · 85%Xiaoxin Chen →
“VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following”
