repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
Atomic-man007/Awesome_Multimodel_LLM
Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Models (MLLM). It covers datasets, tuning techniques, in-context learning, visual reasoning, foundational models, and more. Stay updated with the latest advancement.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation →
- PossiblePossibly related (embedding) · 59%TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference →
- PossiblePossibly related (embedding) · 58%Identifying Interactions at Scale for LLMs →
- PossiblePossibly related (embedding) · 57%Understanding Large Language Models →
- PossiblePossibly related (embedding) · 57%deepseek-ai/DeepSeek-V3 →
- PossiblePossibly related (embedding) · 53%FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence →
- PossiblePossibly related (embedding) · 49%swiss-ai/Apertus-v1.5 70B/8B →
- PossiblePossibly related (embedding) · 51%Setting Up Your Own Large Language Model - Towards Data Science →
Implements
Covers
Related to
Covers (incoming)
newsFLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligencenewsswiss-ai/Apertus-v1.5 70B/8BnewsSetting Up Your Own Large Language Model - Towards Data SciencenewsI developed a 270 million parameter language model entirely from scratch as an independent research projectnewsThe Large Language Model (LLM) understands the world through human texts. A world model is needed f.. - 매일경제newsHow should I approach training this specific ML model for my startup project [D]newsLarge Language Models in Life Science Research: What Scientists Need to Know - Technology NetworksnewsJ-Wash: A novel way to brainwash and customize large language models based on Anthropic's Jacobian-Lens!newsThinking Machines open sources first multimodal language model, Inkling, focused on low cost and 'resistance to censorship' - VentureBeatnewsSeeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]
Implements (incoming)
paperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionpaperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSpaperHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMspaperSwitch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement LearningpaperSigLIP-HD by Fine-to-Coarse SupervisionpaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperA Sovereign, Open-Source Foundation Model for German and EnglishpaperActor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPOpaperExtending LLM Context via Associative Recurrent MemorypaperEvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalpaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperSegregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment AnalysispaperDo We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
Related across the graph
paperHoloCount: A Holistic Visual Counting Benchmark for MLLMspaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationnewsSetting Up Your Own Large Language Model - Towards Data SciencepaperHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMspaperVendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image DetectionnewsThe Large Language Model (LLM) understands the world through human texts. A world model is needed f.. - 매일경제modeldeepseek-ai/DeepSeek-V3newsThinking Machines open sources first multimodal language model, Inkling, focused on low cost and 'resistance to censorship' - VentureBeatnewsHow should I approach training this specific ML model for my startup project [D]paperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioningnewsswiss-ai/Apertus-v1.5 70B/8BpaperExtending LLM Context via Associative Recurrent MemorypaperWordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTSpaperSegregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment AnalysisnewsI developed a 270 million parameter language model entirely from scratch as an independent research projectnewsSeeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]paperCoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty EstimationpaperEvaluating and Understanding Model Editing for Medical Vision Language ModelspaperA Sovereign, Open-Source Foundation Model for German and EnglishpaperSigLIP-HD by Fine-to-Coarse SupervisionpaperDo We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?paperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperUnderstanding Large Language ModelspaperActor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPOnewsLarge Language Models in Life Science Research: What Scientists Need to Know - Technology NetworksnewsFLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual IntelligencenewsJ-Wash: A novel way to brainwash and customize large language models based on Anthropic's Jacobian-Lens!paperSwitch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement LearningpaperEvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalnewsIdentifying Interactions at Scale for LLMspaperTOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
