Read original ↗
paperarXivTrust 82 · PrimaryPublished 8d agoLive · 7d ago

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-level cues from the available audio and visual modalities via pre-trained models. We therefore assemble a wide variety of pre-trained models to extract features that encode mouth movements, face parsing, facial expressions, head pose, gaze tracking, heart rate, audio emotion and speech activity. We further integrate both unimodal and multimodal cue

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Vlad Hondru

    DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

  • LinkedLinked via arxiv author · 85%Florinel Alin Croitoru

    DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

  • LinkedLinked via arxiv author · 85%Iuliana Georgescu

    DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

  • LinkedLinked via arxiv author · 85%A. Sophia Koepke

    DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

  • LinkedLinked via arxiv author · 85%Radu Tudor Ionescu

    DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

authored (incoming)

Related across the graph

Topics