jargon

Applied AI·The API layer

the twenty-second wait becomes an answer that visibly starts in under a second, without the model getting any faster.

Streaming

Draft summary, pending review

Receiving the response token by token as it is generated instead of waiting for the whole thing. Essential for chat UX: it converts a 20 second wait into a response that visibly starts in under a second.

Commonly confused with