When user sends request to a LLM the system needs to set aside memory for the KV cache of that request, but the system does not know in advance how long each response will be

This causes problems because system allocated too much memory in an attempt to be safe

Paged Attention manages KV Cache more efficiently by breaking it into small fixed size blocks called pages, pages don’t have to be next to one another in memory