Read original ↗
newsReddit r/devopsTrust 52 · CommunityPublished 29d agoLive · 28d ago

Cost attribution keeps finding waste that monitoring dashboards miss - how do others handle the gap?

I work on cost attribution for a large internal cloud fleet, and a pattern keeps repeating: our observability stack says everything is healthy, and the invoice says otherwise. Recent example: a batch worker that was fully green in monitoring — no errors, no alerts, normal resource graphs - but was costing ~$4,200/month sitting mostly idle because it was provisioned for a peak workload that moved to a different pipeline two quarters ago. Nobody's dashboard