Comparison
Data cataloguevsData swamp
Data catalogue
you search for 'revenue', find eleven tables with that word in them, and can see which one is actually used by the finance dashboard.
A searchable inventory of datasets with their schemas, owners, lineage, freshness and usage. Its value is disproportionately in usage statistics: 'this table is queried four hundred times a day and that one twice a year' answers the question a description never will. A catalogue populated by hand decays within months, so the ones that work are the ones fed automatically from the warehouse and the orchestrator.
Full entry →Data swamp
there are nine thousand tables in the bucket, four of them are used, and nobody knows which four.
A lake that accumulated data without ownership, documentation or lifecycle, so nothing in it can be trusted or safely deleted. It is the predictable end state of storing everything with no catalogue and no retention, and it is expensive twice: in storage and in every hour spent working out which of five similar tables is the real one. It is worth naming because the remedy is governance and deletion rather than more storage.
Full entry →