Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient GPUs, but operators currently lack the tools to do so without exhaustively profiling each combination. While some predictive models exist, they still require profiling data and struggle to generalize to hardware unseen during training. To address this, we introduce \textit{WattGPU}, featuring two predictive models for mean GPU power draw and Inter-Token Latency (ITL). Our approach leverages only pu

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Implements

Implements (incoming)

authored (incoming)

Covers (incoming)

Related across the graph

repoAMD-AGI/Magpierepokalavai-net/kalavai-clientrepouccl-project/ucclrepolast9/gpu-telemetrynewsIn December 2025, startup Starcloud trained the first large language model ever trained in orbit, using an NVIDIA H100 — the same class of GPU built for Earth's AI data centres, now running roughly 500 kilometres above the planet. - ScienceBlog.comrepogpustack/gpustackpersonMarta Patiño-MartíneznewsGPUHedge: Hedging serverless GPU providers improves cold start p95 latency from 117s to 30s [P]newsNASA Puts Google’s Gemma Large Language Model in Orbitreponovitalabs/pegaflowrepojdermody/brightwirerepogiannisanni/pulsarnews[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer DevicesrepoInfernetProtocol/infernet-protocolnewstried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s)repothejollydev/bezaforge-infrastructurenewsHow NVIDIA’s Inference Software Stack Powers the Lowest Token Costrepoopenlake-project/openlakenewsWe'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D]personMauricio Fadel ArgerichnewsUnderstanding GPU Inference Workloads [D]repoDaoyuanLi2816/llm-gpu-labrepozwmaronek/Beyond-Early-ExitpersonJonathan Fürstrepobeam-cloud/beta9repoxorbitsai/inferencenewsShow HN: Computable – Buy, sell, and redeem GPU for the exact weeks you wantrepoglab-forks/nvidia/TensorRT-LLMnewsHardware startup unveils inference acceleratorrepokekzl/impnewsWattGPU Predicts LLM Inference Power Without Profiling - Let's Data Sciencerepomosecorg/mosec

Topics