SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs
Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of Natural Language Processing (NLP) tasks, but their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Existing approaches to model compression and optimization often rely on coarse-grained pruning or quantization, which can compromise accuracy or require re-training and fine-tuning. In this work, we introduce SelectInfer, a neuron-level optimization framework that enables efficient LLM inference on edge devices through selective neuron loa
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 26%alibaba/MNN →
“Possibly related via embedding similarity 0.57 (not asserted). Timestamp check: artifact slightly before paper (-14d).”
- PossiblePossibly related (embedding) · 53%New Server Hopes to Break Through AI’s “Memory Wall” →
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Huzaifa Shaaban Kabakibo →
“SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs”
- LinkedLinked via arxiv author · 85%Eric Schniedermeyer →
“SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs”
- LinkedLinked via arxiv author · 85%Artem Burchanow →
“SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs”
- LinkedLinked via arxiv author · 85%Zhilin Wang →
“SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs”
