MHA / MQA / GQA

Appears in 1 tutorial

Multi-Head / Multi-Query / Grouped-Query Attention: design choices trading KV-cache size against quality.

As used in LLM Infrastructure →

Multi-Head / Multi-Query / Grouped-Query Attention: design choices trading KV-cache size against quality. GQA (a middle ground) is the modern default for efficient KV cache.