Comparison
Pipeline backfillvsReplay
Pipeline backfill
the logic was wrong for three months, so you re-run the same job for ninety past days and hope nothing downstream notices mid-way.
Re-running a pipeline over historical periods to correct or populate them. Whether it is routine or terrifying is decided entirely by idempotency and by blast radius: a backfill of an idempotent partitioned model is boring, and a backfill of an append-only table is a duplication incident. The other half is resource contention — ninety runs at once will happily starve the scheduled work — which is why concurrency limits and a deliberate order matter more than the code does.
Full entry →Replay
you fixed the bug and reset the consumer back to last Tuesday so it processes the last week of events again correctly.
Re-reading retained messages from an earlier position to rebuild state or recover from a processing bug. It is one of the main reasons to choose a log over a queue. It also re-triggers every side effect the consumer performs, so a consumer that sends emails needs a replay mode before you ever need to replay it.
Full entry →