SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per second (TPS) the primary goal and relaxing--not eliminating--per-token latency requirements; and (2) requests share much of their KV\$-reuse exceeds 80% of request tokens in a production trace from BAILIAN, versus 54-62% in chat. This paper first contributes a systematic study of request scheduling for agents on two real
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%mjason/long →
- PossiblePossibly related (embedding) · 53%Mesh-LLM/mesh-llm →
- PossiblePossibly related (embedding) · 52%tokentopapp/tokentop →
- PossiblePossibly related (embedding) · 52%antoinezambelli/forge →
- PossiblePossibly related (embedding) · 52%harvard-cns/orla →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “SMetric: Rethink LLM Scheduling for Serving Agents with Bala” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “SMetric: Rethink LLM Scheduling for Serving Agents with Bala” ≈ “Unity-Technologies/ml-agents””
- FuzzySimilar title/name (fuzzy) · 84%tensorflow/serving →
“Fuzzy title match (0.92): “SMetric: Rethink LLM Scheduling for Serving Agents with Bala” ≈ “tensorflow/serving””
