Comparison
Canonical datasetvsSystem of record
Canonical dataset
there is one table everyone agrees is the answer for orders, and the other four are explicitly labelled as derived or deprecated.
One agreed authoritative dataset for a given concept, with everything else derived from it. It is easy to say and organisationally difficult, because the competing copies usually exist for good local reasons and removing one costs somebody their workflow. The achievable version is not one table but one lineage root plus a clear statement of which derivations are sanctioned.
Full entry →System of record
two systems disagree about the customer's address and you need to be able to say, without arguing, which one is allowed to be right.
The system whose copy of a fact is authoritative, by agreement rather than by accident. Naming it is what stops a data integration turning into a negotiation every time two copies drift. It matters most where data flows both ways: without an agreed system of record, a sync loop will happily overwrite the correct value with the stale one and neither team will be able to prove which happened.
Full entry →