Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · 2d ago

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint video-audio model. Large-scale T2VA distillation is challenged by modality-imbalanced optimization, the difficulty of continuous-time consistency training at scale, and the quality--diversity trade-off. TurboT2VA addresses these issues with per-modality normalization and a progre

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%openai/whisper-large-v3-turbo

    Fuzzy title match (0.73): “TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation v” ≈ “openai/whisper-large-v3-turbo”

  • FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo

    Fuzzy title match (0.73): “TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation v” ≈ “Tongyi-MAI/Z-Image-Turbo”

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%deepspeedai/DeepSpeed

    Shared author/contributor keys: lai

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • FuzzySimilar title/name (fuzzy) · 59%drumih/turbo-fieldfare

    Fuzzy title match (0.73): “TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation v” ≈ “drumih/turbo-fieldfare”

  • LinkedLinked via arxiv author · 85%Xiaoda Yang

    TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics