Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Unde
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Large Tabular Models Excel Where LLMs Fail →
- PossiblePossibly related (embedding) · 46%The Hidden Infrastructure Challenge Behind Every AI-Generated Avatar - SD Times →
- FuzzySimilar title/name (fuzzy) · 84%huggingface/datasets →
“Fuzzy title match (0.92): “Computational Humor with Multimodal LLMs: Methods, Datasets,” ≈ “huggingface/datasets””
- FuzzyOverlapping authors or contributors · 62%rasbt/LLMs-from-scratch →
“Shared author/contributor keys: yin”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- LinkedLinked via arxiv author · 85%Tuo Liang →
“Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges”
- LinkedLinked via arxiv author · 85%Haozhe Huang →
“Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges”
- LinkedLinked via arxiv author · 85%Disheng Liu →
“Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges”
