repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 3d ago
modelscope/mcore-bridge
MCore-Bridge: Providing Megatron-Core model definitions for state-of-the-art large models and making Megatron training as simple as Transformers — with support for 300+ large language models (Qwen3-Next, GLM-5.2, Deepseek-V4, MiniMax-2.7, ...) and 200+ multimodal large models (Qwen3.5, Qwen3-Omni, Gemma4, ...).
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%deepseek-ai/Janus-Pro-7B →
- PossiblePossibly related (embedding) · 49%openai/whisper-large-v3 →
- PossiblePossibly related (embedding) · 48%deepseek-ai/DeepSeek-V3 →
- PossiblePossibly related (embedding) · 48%H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P] →
- PossiblePossibly related (embedding) · 47%Introducing Gemma 4 12B: a unified, encoder-free multimodal model →
- PossiblePossibly related (embedding) · 51%Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost - MarkTechPost →
- PossiblePossibly related (embedding) · 49%With all the Kimi drama I feel like I want to download all the current best models in case there is a ridiculous knee jerk political move pulled →
- PossiblePossibly related (embedding) · 46%Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork! →
Related to
Covers
Covers (incoming)
newsKimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost - MarkTechPostnewsWith all the Kimi drama I feel like I want to download all the current best models in case there is a ridiculous knee jerk political move pullednewsUncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!
Implements (incoming)
Related across the graph
modeldeepseek-ai/DeepSeek-V3newsWith all the Kimi drama I feel like I want to download all the current best models in case there is a ridiculous knee jerk political move pullednewsKimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost - MarkTechPostmodeldeepseek-ai/Janus-Pro-7BpaperDo We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?modelopenai/whisper-large-v3newsH64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]newsUncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!newsIntroducing Gemma 4 12B: a unified, encoder-free multimodal model
