Platform & DevOps·Running it: capacity, failure and recovery
the plan for the day a whole region, account or dataset is gone, written down and hopefully rehearsed once a year.
Disaster recovery
Also calledDR, business continuity
The prepared response to losing an entire environment — a region, an account, a dataset — as opposed to the ordinary failures your architecture absorbs. It is defined by two numbers, how much time and how much data you can lose, and everything else follows from them. Its recurring failure is that the plan exists and has never been executed, so nobody knows that the restore takes eleven hours or that the credentials to run it are in the environment that is gone.