Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, imag
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 84%Thysrael/Horizon →
“Fuzzy title match (0.92): “Long-Horizon Audio-Visual Generation for Persistent Stories ” ≈ “Thysrael/Horizon””
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Nan Duan →
“Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds”
- LinkedLinked via arxiv author · 85%Haoyang Huang →
“Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds”
- LinkedLinked via arxiv author · 85%Weiyang Jin →
“Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds”
