jargon

Applied AI·Evaluation and reliability

nothing in your code changed but the outputs are worse than they were in March, because the traffic or the model underneath moved.

Drift

Draft summary, pending review

Quality change over time without any deploy on your side: providers update models, user inputs shift, corpora age. Detected by scheduled eval runs and monitored online metrics; pinned model versions reduce it but expire when versions retire.

Commonly confused with