KV Cache Cost Attribution: Who's Actually Paying for GPU Memory
KV cache cost attribution means tying the memory side of your inference bill to the requests, sessions. And tenants that actually created it.
That's the whole idea. And almost nobody does it, because the bill you get doesn't show memory at all. It shows tokens.
The