repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
tokentopapp/tokentop
htop for your AI costs — real-time terminal monitoring of LLM token usage and spending across providers and coding agents
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%TraceLab: Characterizing Coding Agent Workloads for LLM Serving →
- PossiblePossibly related (embedding) · 55%I spent ~4.5 months building a free, self-hosted AI gateway: one endpoint for 237 providers (90+ free), auto-fallback, and a token-compression pipeline (MIT) →
- PossiblePossibly related (embedding) · 52%NVIDIA Unlocks AI Compute at Scale, Inviting Partners to Power the AI Infrastructure Buildout →
- PossiblePossibly related (embedding) · 52%How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost →
- PossiblePossibly related (embedding) · 52%NVIDIA Unlocks AI Compute at Scale, Inviting Capital Partners to Power the AI Infrastructure Buildout →
- PossiblePossibly related (embedding) · 52%SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling →
- PossiblePossibly related (embedding) · 49%Unlimited AI tokens aren't unlimited after all as US Army burns through supply →
- PossiblePossibly related (embedding) · 55%Show HN: Frugal Tokens – explore costs and usage across coding agents →
Implements
Covers
newsI spent ~4.5 months building a free, self-hosted AI gateway: one endpoint for 237 providers (90+ free), auto-fallback, and a token-compression pipeline (MIT)newsNVIDIA Unlocks AI Compute at Scale, Inviting Partners to Power the AI Infrastructure BuildoutnewsHow NVIDIA’s Inference Software Stack Powers the Lowest Token CostnewsNVIDIA Unlocks AI Compute at Scale, Inviting Capital Partners to Power the AI Infrastructure Buildout
Implements (incoming)
Covers (incoming)
newsUnlimited AI tokens aren't unlimited after all as US Army burns through supplynewsShow HN: Frugal Tokens – explore costs and usage across coding agentsnewsNvidia shows off Vera Rubin platform for tokenmaxxingnewsThe price is wrong: AI cost calculation has to consider task completion rates, not just token costsnewsAnthropic's extravagant tokenizer complicates AI pricingnewsHow Much Does It Actually Cost to Run a Local LLM? (Euros per Million Tokens, Measured) - Towards Data SciencenewsHow Google’s New Gemini Rates Work and How to Track Your Usage
Related across the graph
paperTraceLab: Characterizing Coding Agent Workloads for LLM ServingnewsAnthropic's extravagant tokenizer complicates AI pricingpaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingnewsShow HN: Frugal Tokens – explore costs and usage across coding agentsnewsNVIDIA Unlocks AI Compute at Scale, Inviting Capital Partners to Power the AI Infrastructure BuildoutnewsHow Google’s New Gemini Rates Work and How to Track Your UsagenewsHow NVIDIA’s Inference Software Stack Powers the Lowest Token CostnewsNvidia shows off Vera Rubin platform for tokenmaxxingnewsUnlimited AI tokens aren't unlimited after all as US Army burns through supplynewsHow Much Does It Actually Cost to Run a Local LLM? (Euros per Million Tokens, Measured) - Towards Data SciencenewsI spent ~4.5 months building a free, self-hosted AI gateway: one endpoint for 237 providers (90+ free), auto-fallback, and a token-compression pipeline (MIT)newsNVIDIA Unlocks AI Compute at Scale, Inviting Partners to Power the AI Infrastructure BuildoutnewsThe price is wrong: AI cost calculation has to consider task completion rates, not just token costs
