newsHugging FaceTrust 88 · LabPublished 7d agoLive · 5d ago
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Vision as Unified Multimodal Generation →
- PossiblePossibly related (embedding) · 54%ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models →
- PossiblePossibly related (embedding) · 54%See & Sniff: Learning Visuo-Olfactory Representations →
- PossiblePossibly related (embedding) · 53%MedUAG: Unified Understanding and Generation for Medical Multimodal Models →
- PossiblePossibly related (embedding) · 53%VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression →
- PossiblePossibly related (embedding) · 48%MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation →
- PossiblePossibly related (embedding) · 53%PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond Decoding →
- PossiblePossibly related (embedding) · 61%SenseNova-U1.5: Towards Native Unified Visual Intelligence →
Covers
paperVision as Unified Multimodal GenerationpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperSee & Sniff: Learning Visuo-Olfactory RepresentationspaperMedUAG: Unified Understanding and Generation for Medical Multimodal ModelspaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
Covers (incoming)
Related across the graph
paperMedUAG: Unified Understanding and Generation for Medical Multimodal ModelspaperPIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond DecodingpaperMEOX: Compact Multimodal Mixture-of-Experts for Earth ObservationpaperENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelspaperSenseNova-U1.5: Towards Native Unified Visual IntelligencepaperSee & Sniff: Learning Visuo-Olfactory RepresentationspaperVision as Unified Multimodal GenerationpaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
