newsReddit r/LocalLLaMATrust 52 · CommunityPublished 27d agoLive · 27d ago
[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%NexusGPU/tensor-fusion →
- PossiblePossibly related (embedding) · 55%zwmaronek/Beyond-Early-Exit →
- PossiblePossibly related (embedding) · 54%DaoyuanLi2816/llm-gpu-lab →
- PossiblePossibly related (embedding) · 54%WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs →
- PossiblePossibly related (embedding) · 53%pytorch/pytorch →
- PossiblePossibly related (embedding) · 54%huggingface/optimum-intel →
- PossiblePossibly related (embedding) · 54%bassrehab/triton-kernels →
- PossiblePossibly related (embedding) · 55%arizqi/cpubrrr →
