The KV Cache: Sizing It, and the Attention Variants That Shrink It

What the KV cache costs per token, and how MQA, GQA and MLA rewrite attention to make that number smaller.

August 12, 2026 · 10 min · Sandeep Kumar