STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic consistency. A model may describe the intended target correctly while generating an inconsistent edit. This exposes an understanding-generation alignment gap: linguistic and visual outputs live in different spaces, yet should be governed by the same target semantics. We study this gap in image editing, where an instruction defines a target state that can be both described and visually realized. Given a source image and
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Ye Wang →
“STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs”
- LinkedLinked via arxiv author · 85%Hongjun Wang →
“STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs”
- LinkedLinked via arxiv author · 85%Minghao Fang →
“STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs”
- LinkedLinked via arxiv author · 85%Tongyuan Bai →
“STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs”
