PagedAttention

How PagedAttention uses virtual-memory-style KV-cache paging to reduce fragmentation and increase LLM serving throughput.

March 3, 2026 · 5 min · Sandeep Kumar