TL

Tag

Attention

2 posts· All tags

Inference13 min

Attention, and How to Shrink It

The KV cache is num_layers × num_heads × 2 × seq_len × d_head × bytes, and every way to shrink attention attacks one of those terms. MQA and GQA share keys across heads, MLA compresses them to a latent that beats full attention on quality, FlashAttention computes the exact same thing with a fraction of the memory traffic, and the frontier goes after the rest — layers, sparsity, and the quadratic itself.