Applied AI·Tokens
you time it yourself on your own hardware, because the number in the announcement was measured on something else.
Tokens per second
Draft summary, pending review
The standard speed metric for generation. Hosted APIs commonly stream tens to low hundreds of tokens per second; local inference speed depends on hardware, model size and quantisation. Measure it yourself rather than trusting marketing numbers.