Adaptive Inference Batching using Policy Gradients
Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic. We investigate whether reinforcement learning (RL) can learn adaptive batching and routing policies that outperform these heuristics, training REINFORCE and PPO agents on a discrete-event simulator validated against queuing theory and production traces (Azure Functions, BurstGPT). We formulate the problem as an MDP over queue state, request type and GPU availability, evalu
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Scaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway Deployment →
- PossiblePossibly related (embedding) · 46%inference-gateway/inference-gateway →
- PossiblePossibly related (embedding) · 48%AgileRL/AgileRL →
- FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference →
“Fuzzy title match (0.92): “Adaptive Inference Batching using Policy Gradients” ≈ “xorbitsai/inference””
- LinkedLinked via arxiv author · 85%Ruslan Sharifullin →
“Adaptive Inference Batching using Policy Gradients”
