Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens. These tokens are derived from a codebook that maps embeddings to quantized visual patterns. The language-like architecture enables unified multimodal models to effectively capture text conditional information for generation, making them promising for text-to-image tasks. This also raises an interesting question: how safe are the images generated in such an autoregressive way? In this work, we propose iterative self
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownDiffusionGemma: 4x faster text generation →
- LinkedLinked via unknownDiffuse-XL →
- LinkedLinked via unknownIntroducing Gemma 4 12B: a unified, encoder-free multimodal model →
- PossiblePossibly related (embedding) · 47%I built my 'first' flow matching image generator, here's what I learned [P] →
- PossiblePossibly related (embedding) · 51%Autoencoders: Learning Through Reconstruction - Snowflake →
