Applied AI·Prompting
you change a word, re-run the eval set, look at what broke, and change it back, instead of reading one good output and declaring victory.
Prompt iteration
Draft summary, pending review
The real workflow: write, run against an eval set, inspect failures, revise, repeat. Without the eval set (section 8) this degenerates into superstition, changing words and hoping.