The KV Cache: Sizing It, and the Attention Variants That Shrink It
What the KV cache costs per token, and how MQA, GQA and MLA rewrite attention to make that number smaller.
What the KV cache costs per token, and how MQA, GQA and MLA rewrite attention to make that number smaller.