Cost per token

Appears in 1 tutorial

GPU $/hour ÷ tokens/hour at the operating point; the metric leadership cares about.

As used in LLM Infrastructure →

GPU $/hour ÷ tokens/hour at the operating point; the metric leadership cares about. Optimizations lower it by raising tokens/hour per GPU.