TL

Tag

Continuous batching

1 post· All tags

Inference11 min

Scheduling: How Continuous Batching and Paged Attention Fill a GPU

Orca reported 36.9× the throughput of FasterTransformer at the same latency without touching the model's math; a year later vLLM's profiling found as little as a fifth of KV-cache memory holding actual tokens. Two papers, two mechanisms — re-form the batch every iteration, page the cache in 16-token blocks — and everything else an LLM scheduler does today is policy.