Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-2-7b-chat-hf →
“Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-2-7b-chat-hf””
- FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-2-7b →
“Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-2-7b””
- FuzzySimilar title/name (fuzzy) · 59%meta-llama/Meta-Llama-3-8B →
“Fuzzy title match (0.73): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Meta-Llama-3-8B””
- PossiblePossibly related (embedding) · 27%vllm-project/vllm →
“Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact slightly before paper (-50d).”
- PossiblePossibly related (embedding) · 25%alibaba/MNN →
“Possibly related via embedding similarity 0.55 (not asserted). Timestamp check: artifact slightly before paper (-46d).”
- FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-3.1-8B-Instruct →
“Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-3.1-8B-Instruct””
- FuzzySimilar title/name (fuzzy) · 87%meta-llama/Llama-3.3-70B-Instruct →
“Fuzzy title match (0.94): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “meta-llama/Llama-3.3-70B-Instruct””
- FuzzySimilar title/name (fuzzy) · 59%run-llama/llama_index →
“Fuzzy title match (0.73): “Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs” ≈ “run-llama/llama_index””
