jargon

Applied AI·Evaluation and reliability

the incident came in and you had every prompt, output, model version, latency and cost to look at instead of guesses.

Observability

Draft summary, pending review

Logging, tracing and metrics for LLM features: every call's prompt, output, tokens, latency, cost, model version and outcome. Without it, production incidents reduce to guessing. It is the same discipline you apply to services, with prompts and completions as the payload.

Commonly confused with