jargon

Applied AI·The API layer

you set it too low and the answer got cut off mid-sentence, then set it too high and a runaway response ate the budget.

max_tokens

Draft summary, pending review

The ceiling on output length for one response. Set it deliberately: too low truncates answers mid-sentence, too high lets a runaway response burn money. It is a budget, not a target; the model does not aim for it.