ChatImage: Navigating Long-Form LLM Answers through Interactive Images
Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which makes fine-grained inspection, navigation, and return visits difficult. We present ChatImage, a system that converts long-form LLM answers into interactive visual images. Given a textual answer, ChatImage first normalizes its content into structured visual modules, plans a visual layout, and renders a coherent image. It then applies a second grounding pass to the rendered image with vision grounding models such as LocateAnything and MiMo-Vision, wi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%VioletVision-3B →
- PossiblePossibly related (embedding) · 53%llmsresearch/llm-flashcards →
- PossiblePossibly related (embedding) · 49%SethRobinson/aitools_client →
- PossiblePossibly related (embedding) · 49%icarito/gtk-llm-chat →
- PossiblePossibly related (embedding) · 48%Atomic-man007/Awesome_Multimodel_LLM →
- FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo →
“Fuzzy title match (0.73): “ChatImage: Navigating Long-Form LLM Answers through Interact” ≈ “Tongyi-MAI/Z-Image-Turbo””
- LinkedLinked via arxiv author · 85%Wencan Jiang →
“ChatImage: Navigating Long-Form LLM Answers through Interactive Images”
- LinkedLinked via arxiv author · 85%Jiangning Zhang →
“ChatImage: Navigating Long-Form LLM Answers through Interactive Images”
