jargon

Applied AI·Production patterns and cost

you allocate the window in code, so much for retrieved context and so much for the answer, instead of discovering the ceiling at 2am.

Token budgeting

Draft summary, pending review

Explicit token allocation per request: so much for system prompt, so much for retrieved context, history and output. Enforced in code with counting and truncation rules, it is what prevents context overflow incidents at 2am.