Attention head

Appears in 1 tutorial

One of several parallel attention computations; more heads = richer attention, bigger KV cache.

As used in LLM Infrastructure →

One of several parallel attention computations; more heads = richer attention, bigger KV cache.