Metadata Management: Inside NDMO's Data Catalog & Metadata Domain
8 min read · June 7, 2026
Of the 15 domains in the NDMO data management framework, Data Catalog & Metadata is the one teams most often underestimate. The name sounds like a documentation chore — a librarian's task sitting next to weighty domains like data governance and personal data protection. In practice it is the opposite. This is the domain every other domain writes into: classification labels, ownership records, quality scores, and retention rules are all metadata, and all of them need a reliable place to live. This article unpacks what metadata actually is, what the domain expects from you in practice, why it underpins the other fourteen domains, and how organizations mature from tribal knowledge to a living catalog.
What metadata actually is
Metadata is data that describes data. It answers the questions that arrive before any analysis can begin: What is this dataset? What does it mean? Who owns it? How sensitive is it? When was it last updated? Practitioners split it into three types, and the distinction matters because each type has a different source, a different audience, and a different maintenance burden.
Business metadata
Business metadata captures what the data means to the organization:
- Definitions in business language — what exactly counts as an "active customer" or a "completed transaction"
- Links to agreed terms in the business glossary
- Ownership and stewardship assignments — who answers for this data
- Classification labels indicating sensitivity
- Regulatory flags, such as whether a dataset contains personal data governed by the PDPL
Business metadata is the hardest type to produce, because no scanner can extract it from a database. It comes from people, through agreement.
Technical metadata
Technical metadata describes how the data is physically represented:
- Schemas, tables, columns, and data types
- Where the data lives — which database, which system, which environment
- Formats, constraints, and keys
- The interfaces and APIs that expose it
Technical metadata is the easiest to collect, because it already exists inside source systems and can be harvested automatically.
Operational metadata
Operational metadata records what happens to the data as it moves and gets used:
- Pipeline runs, load times, and freshness
- Data quality measurement results
- Access patterns and usage statistics
- Lineage events — what feeds this table, and what it feeds in turn
A concrete example ties the three together. Take a "Customers" table. Technical metadata tells you the national ID column is a ten-character text field in the CRM database. Business metadata tells you it holds the national identity number, is owned by the customer care department, is classified as confidential, and is personal data under the PDPL. Operational metadata tells you the table refreshed at 2 a.m., passed fourteen of fifteen quality checks, and feeds seven downstream reports. Each type answers a different question — and both an assessor and a working analyst eventually need all three.
What the NDMO domain expects in practice
Like every domain in the framework, this one decomposes into controls and then into auditable specifications. Read as a whole, the expectations come down to five things:
- A complete inventory of data assets. A register of the entity's datasets, databases, and reports — not a one-off list, but a maintained record whose coverage visibly grows over time.
- Business definitions, not just technical listings. A business glossary of agreed terms, linked to the physical assets that implement them, so that "beneficiary" means the same thing in every report.
- A metadata standard. An agreed minimum set of attributes that every registered asset must carry: name, description, owner, steward, classification, source system, update frequency.
- Ownership and stewardship of the metadata itself. Metadata is an asset too. Named stewards must answer for its accuracy and completeness, with review cycles to prove it.
- Processes that keep it current. New systems get registered, schema changes get reflected, and harvesting is automated wherever the sources allow it.
When assessment time comes, the evidence is concrete: the register itself, its coverage percentage, how many assets carry definitions and owners, and proof that the catalog has been updated since it was created. A catalog exported once before the assessment and never touched again scores like exactly what it is.
The foundation the other fourteen domains build on
Stated plainly: most NDMO domains produce metadata as their evidence and consume metadata as their input.
- Data classification attaches sensitivity labels to assets in the catalog — you cannot classify an inventory you do not have.
- Data quality defines rules against fields the metadata describes, and writes its scores back as operational metadata.
- Personal data protection starts with knowing where personal data lives — which is a metadata question before it is a legal one.
- Data sharing and open data cannot publish or exchange a dataset that has no description, no owner, and no classification.
- Data lineage is itself metadata: a record of how data moves between systems.
- Data governance records its central artifacts — ownership and stewardship assignments — as metadata attached to assets.
This is why postponing the domain is expensive. Entities that defer it discover that classification, quality, and protection work all stall, waiting for an inventory that should have existed first.
Maturity stages: from tribal knowledge to active metadata
Metadata capability tends to develop through five recognizable stages:
- Tribal knowledge. Metadata lives in the heads of long-serving employees and in scattered files. Answering "what does this column mean?" means finding the right person — and hoping they have not resigned.
- Static documentation. Spreadsheets and documents describe the main systems. Better than nothing, but the content starts decaying the day it is written, and within a year nobody trusts it.
- Centralized catalog. A single register with a defined metadata standard, named stewards, and manual curation. A major step — assessable, auditable, but still labor-intensive.
- Automated and connected. The catalog harvests technical metadata directly from source systems; schema changes appear automatically; quality scores and lineage flow in on their own. Human effort shifts to business metadata, where people genuinely add value.
- Active metadata. Metadata drives action rather than just describing it: classification labels trigger access policies, schema changes alert downstream owners, and business users search the catalog daily as a habit, not a mandate.
Most organizations entering their first assessment sit at stage one or two. A realistic ambition is stage three within the first compliance cycle, and stage four in the one after. Attempting to jump straight to stage five usually produces an expensive platform nobody opens.
A practical first quarter
If you are starting from scratch, a focused ninety days looks like this:
- Scope deliberately. Pick the three to five most critical systems — not everything. Coverage of what matters beats thin coverage of everything.
- Agree a minimum standard. Ten to fifteen attributes per asset, no more. A standard nobody can fill in is a standard nobody follows.
- Inventory and assign stewards. Register the in-scope assets and name a steward for each — a real person with the standing to keep definitions honest.
- Automate the technical layer. Connect the catalog to source systems so schemas, tables, and columns arrive and stay current without manual entry.
- Build the glossary from real disputes. Start with the terms that already cause arguments between departments — "customer," "revenue," "active." Resolving those creates believers.
Done in this order, the catalog stops being a compliance artifact and becomes the working memory of the organization — which is what the domain intended all along.
This is also the one domain where tooling genuinely changes the economics: a manually maintained register decays between assessments, while automated harvesting keeps it alive by default. If you are evaluating platforms, Goava's data catalog collects technical metadata automatically and keeps business definitions, ownership, and classification in one place — mapped directly to the NDMO evidence you will be asked to show.
Monthly digest
Monthly data governance insights for organizations operating in Saudi Arabia.