jargon

Applied AI·Evaluation and reliability

you serve the new prompt to half the traffic and the old one to the other half, and let the outcome metric decide instead of your opinion.

A/B testing

Draft summary, pending review

Serving variants (prompts, models, pipelines) to split traffic and comparing outcome metrics. The same discipline as any product experiment, applied to probabilistic components; your existing statistics instincts transfer unchanged.

Commonly confused with