Applied AI·Evaluation and reliability
the new prompt goes to two percent of traffic first, and you watch quality and cost before letting the rest through.
Canary deployment
Draft summary, pending review
Rolling a prompt or model change to a small traffic slice first, watching quality and cost metrics, then widening. Treat prompts as deployable artefacts with the same ceremony as code, because they break things the same way.