Applied AI·The API layer
you set it too low and the answer got cut off mid-sentence, then set it too high and a runaway response ate the budget.
max_tokens
Draft summary, pending review
The ceiling on output length for one response. Set it deliberately: too low truncates answers mid-sentence, too high lets a runaway response burn money. It is a budget, not a target; the model does not aim for it.