Inference12 min
Attention, from First Principles
Attention is the one operation that lets a token look back at the whole sequence — and it is also where the KV cache comes from. This post derives scaled dot-product attention from the ground up, then watches the cache, and the memory wall from the last post, fall straight out of the math. The first of two; the next is about making it cheap.