Read original ↗
paperarXivTrust 82 · PrimaryPublished 11d agoLive · 8d ago

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-2-7b-chat-hf

    Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-2-7b-chat-hf”

  • FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-2-7b

    Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-2-7b”

  • FuzzySimilar title/name (fuzzy) · 59%meta-llama/Meta-Llama-3-8B

    Fuzzy title match (0.73): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Meta-Llama-3-8B”

  • PossiblePossibly related (embedding) · 27%vllm-project/vllm

    Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact slightly before paper (-50d).

  • PossiblePossibly related (embedding) · 25%alibaba/MNN

    Possibly related via embedding similarity 0.55 (not asserted). Timestamp check: artifact slightly before paper (-46d).

  • FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-3.1-8B-Instruct

    Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-3.1-8B-Instruct”

  • FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-3.3-70B-Instruct

    Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-3.3-70B-Instruct”

  • FuzzySimilar title/name (fuzzy) · 59%run-llama/llama_index

    Fuzzy title match (0.73): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “run-llama/llama_index”

Has model

Related to

Implements (incoming)

authored (incoming)

Related across the graph

Topics