FlashAttention
An exact, memory-efficient attention kernel that avoids writing big intermediates to slow memory; speeds attention, especially for long contexts.
An exact, memory-efficient attention kernel that avoids writing big intermediates to slow memory; speeds attention, especially for long contexts. No quality loss.