Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and c
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%browser-use/browser-use →
“Shared author/contributor keys: lee”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%Kong/kong →
“Shared author/contributor keys: kong”
- LinkedLinked via arxiv author · 85%Sreyan Ghosh →
“Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos”
- LinkedLinked via arxiv author · 85%Arushi Goel →
“Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos”
- LinkedLinked via arxiv author · 85%Kaousheik Jayakumar →
“Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos”
- LinkedLinked via arxiv author · 85%Lasha Koroshinadze →
“Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos”
- LinkedLinked via arxiv author · 85%Nishit Anand →
“Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos”
