Applied AI·Evaluation and reliability
nothing in your code changed but the outputs are worse than they were in March, because the traffic or the model underneath moved.
Drift
Draft summary, pending review
Quality change over time without any deploy on your side: providers update models, user inputs shift, corpora age. Detected by scheduled eval runs and monitored online metrics; pinned model versions reduce it but expire when versions retire.