Comparison
Data ownershipvsTrust threshold
Data ownership
the table is wrong and the question 'who owns this' takes two days and three Slack channels to answer.
A named team accountable for a dataset's correctness, freshness and lifecycle. Without it, quality problems are everyone's to notice and nobody's to fix, which is the default state of any warehouse over a certain size. Ownership recorded in the catalogue rather than in someone's memory is what makes an incident routable at three in the morning.
Full entry →Trust threshold
some tables are marked as certified with tests and an owner, and the rest are clearly labelled as somebody's experiment.
An explicit statement of how much a given dataset can be relied on, published alongside it. It exists because a warehouse always contains a mixture of production models and half-finished exploration, and an analyst cannot tell them apart by looking. Tiering is cheaper than raising everything to production standard and more honest than pretending the difference is not there.
Full entry →