An application that attempts to predict future outcomes through probability estimates is called:
Reactive analytics
Descriptive analytics
Predictive analytics
Dimensional analytics
Just-in-time reporting
Predictive analytics applies statistical, mathematical or machine-learning models to historical and current data in order to estimate future events, behaviours or outcomes. DAMA's treatment of analytics positions predictive analytics as a use of managed data in which patterns discovered from existing information are used to project what is likely to occur rather than merely describe what has already occurred. DMBOK2 explicitly places predictive analytics among advanced uses of data and notes the evolution of Business Intelligence from retrospective analysis toward predictive methods.
Probability estimation is therefore the distinguishing feature in this question. Descriptive analytics summarizes historical conditions—what happened. Predictive analytics estimates what may happen and commonly produces probabilities, scores, classifications or forecasts. “Reactive analytics,†“dimensional analytics,†and “just-in-time reporting†do not describe this statistical forecasting function.
Data Quality is fundamental to predictive performance. Missing values, inaccurate labels, duplicate entities, inconsistent categories, stale observations and biased source populations can materially distort estimated probabilities. Profiling should therefore assess distributions, completeness, outliers and consistency before modelling. Metadata must preserve definitions, lineage and transformation logic, while governed Master and Reference Data provide consistent classifications across training and operational datasets.
Predictive outputs themselves can also be monitored statistically for unexpected distribution shifts and unreasonable patterns.
Reference Topics: DAMA-DMBOK2 — Data Warehousing and Business Intelligence; Predictive Analytics; Chapter 13 — Profiling, Statistical Quality Control, Reasonableness, Accuracy and Completeness.
===============
In 2009, ARMA International published GARP for managing records and information. GARP stands for:
Generally Accepted Recordkeeping Principles
Generally Available Recordkeeping Practices
G20 Approved Recordkeeping Principles
Global Accredited Recordkeeping Principles
Gregarious Archive of Recordkeeping Processes
GARP stands for Generally Accepted Recordkeeping Principles. ARMA International developed the framework to describe the characteristics of an effective records and information-management program. The principles address accountability, transparency, integrity, protection, compliance, availability, retention, and disposition.
These principles align closely with DAMA-DMBOK2's treatment of Document and Content Management. Information must be created, organized, protected, maintained, retrieved, retained, and ultimately disposed of according to business, legal, regulatory, and historical requirements. Records management is therefore not simply about long-term storage. It establishes controls throughout the information lifecycle.
For example, the Integrity principle addresses the authenticity and reliability of records. Availability requires information to be retrievable accurately and efficiently. Retention requires organizations to keep records for justified periods, while Disposition governs appropriate final handling after those obligations expire.
These controls interact directly with Data Quality. Records that cannot be found, trusted, interpreted, or demonstrated to be authentic are not fit for their intended purpose even if they physically exist.
Metadata Management supplies classification, retention, provenance, ownership, and lifecycle metadata necessary to implement these principles consistently.
Reference Topics: DAMA-DMBOK2 — Document and Content Management; Records Management; GARP; Retention; Disposition; Integrity; Metadata Management.
===============
In data modelling practice, entities are linked by:
Indexes
Triggers
Cardinality
Relationships
Processes
In a data model, entities are linked by relationships. An entity represents a distinguishable business concept or thing about which the organization stores information, while a relationship expresses how one entity is associated with another.
For example, a Customer places an Order, an Order contains Order Lines, and a Product appears on an Order Line. Relationships therefore capture business semantics rather than merely physical implementation details. Cardinality is an important characteristic of a relationship because it specifies how many instances of one entity may or must be associated with instances of another; however, cardinality is not itself the general mechanism by which entities are linked. DAMA-oriented modelling material identifies relationships as the correct linkage concept.
Indexes and triggers are physical database mechanisms. Processes describe business activities and may interact with entities, but they do not define entity-to-entity structure within a data model.
The Data Quality implications are significant. Properly defined relationships establish expectations for referential integrity. For example, an Order referencing a nonexistent Customer indicates an integrity defect. Profiling can test orphaned foreign keys, invalid relationship cardinalities, and inconsistent associations.
Relationships and their cardinalities should therefore be documented as metadata and enforced through appropriate database, application, or quality controls.
Reference Topics: DAMA-DMBOK2 Chapter 5 — Entities, Relationships and Cardinality; Logical Data Modeling; Chapter 13 — Integrity and Consistency; Metadata Management.
===============
Three common interaction models for data integration are:
Point to point, wheel and spoke, public and share
Record and pass, copy and send, read and write
Plane to point, harvest and seed, publish and subscribe
Point to point, hub and spoke, publish and subscribe
Straight copy, curved copy, roundabout copy
DAMA-DMBOK2 identifies Point-to-Point, Hub-and-Spoke, and Publish-Subscribe as three common interaction models used to transfer data between systems.
In Point-to-Point, one system communicates directly with another. It is straightforward for a small number of systems but becomes difficult to maintain as interfaces proliferate. In Hub-and-Spoke, systems exchange information through a central hub, reducing direct dependencies and supporting greater consistency. In Publish-Subscribe, producing systems publish data while interested consumers subscribe to the relevant data service or event stream. DMBOK2 presents these three patterns explicitly within Data Integration and Interoperability.
The choice of pattern affects Data Quality as well as architecture. Point-to-point designs can create inconsistent transformations when several applications independently consume the same source. Hub-and-spoke designs can centralize transformations and reference mappings. Publish-subscribe supports scalable distribution but requires controlled message definitions and metadata so that every subscriber interprets data consistently.
DAMA therefore recommends choosing an interaction model according to functional, performance, latency, support, and governance requirements rather than allowing integration patterns to emerge independently.
Reference Topics: DAMA-DMBOK2 Chapter 8 — Data Integration and Interoperability; Interaction Models; Point-to-Point; Hub-and-Spoke; Publish-Subscribe; Canonical Models.
===============
Three source systems provide different telephone numbers for the same customer. An MDM hub selects one number according to approved source-priority rules. This is an example of:
Survivorship
Encryption
Normalization
Archiving
The process is Survivorship. In Master Data Management, survivorship determines which value should become the preferred or authoritative representation when multiple source records contain conflicting values for the same attribute.
Rules may prioritize particular systems, use recency, trust scores, verification status, completeness, or combinations of factors. For example, a verified customer-service update might outrank an older marketing-system telephone number.
Survivorship occurs after or in conjunction with matching and entity resolution. The organization first determines that records from different sources represent the same real-world customer and then determines which attribute values should populate the mastered representation.
The process requires governance because the "best" value is a business decision, not merely a technical one. Data Stewards should approve the rules, and metadata should record source priority, rule logic, and lineage.
DAMA's framework positions Reference and Master Data Management as the discipline responsible for ensuring consistent core entities across the organization.
Reference Topics: DAMA-DMBOK2 Chapter 10 — Master Data Management; Matching; Survivorship; Authoritative Values; Chapter 13 — Consistency and Accuracy.
===============
Which of the following is the best example of a 'documented data quality rule'?
Each transaction data file holding customer transactions must be kept confidential to the authorized users within the operations team
Every transaction recorded must be processed by 12:05 am by authorized personnel who will validate balances and data delivery to branches
All transaction data in the core banking systems need to be processed at 12:05 am each day regardless of the business calendar day and timezone
The transaction data from all satellite systems needs to be ready by 12:05 am in order to feed the overnight batching window, to ensure branches have access to actual customer balances
The transaction data from all satellite systems needs to reflect actual customer balances each morning
Option D is the strongest example because it expresses a Data Quality requirement that is specific, measurable, associated with identified data, and explicitly connected to business need. It identifies the relevant data—transaction data from satellite systems—sets a measurable timeliness threshold of 12:05 a.m., and explains the business reason: ensuring that branches have access to actual customer balances for the required operating window. Published versions of this DAMA question identify the same choice.
This is fundamentally a Timeliness rule. DMBOK2 emphasizes that Data Quality requirements should be defined according to business expectations rather than vague aspirations. A useful rule needs sufficient precision to support objective measurement, exception reporting, and remediation.
Option E captures an important quality expectation—actual customer balances—but it lacks the same precise operational threshold. Option A is a security/confidentiality requirement rather than a Data Quality rule. Options B and C primarily describe processing procedures and do not express the business-quality requirement as cleanly.
Once documented, the rule should be managed as metadata, assigned to an accountable steward, linked to its Critical Data Elements, monitored through metrics, and escalated when the specified threshold is breached.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Data Quality Requirements; Data Quality Rules; Timeliness; Measurement; Business Impact; Metadata Management.
===============
Obfuscation or redaction of data is the practice of:
Reducing the size of large databases
Selling data
Making information anonymous or removing sensitive information
Making information available to the public
Organizing data into meaningful groups
DAMA-DMBOK2 defines obfuscation or redaction as making information anonymous or removing sensitive information. The purpose is to reduce the risk that protected, confidential, or personally identifiable information can be exposed to users or processes that do not require access to the original values.
Obfuscation can involve masking, substitution, shuffling, temporal variation, partial display, or other methods that change what the recipient sees while preserving sufficient utility for the intended activity. For example, a customer-service agent may see only the final digits of an account identifier, or a development team may receive realistic but anonymized production-derived test data.
Redaction may remove information entirely from a representation when there is no legitimate requirement to expose it.
The technique is therefore a Data Security and privacy control, not database compression, publication, or classification.
There is also an important Data Quality consideration. Masked or obfuscated data used for testing must remain structurally valid and retain required relationships so that applications behave realistically. Metadata should indicate that a dataset has been transformed for privacy purposes so consumers do not mistake masked values for authoritative production data.
Reference Topics: DAMA-DMBOK2 Chapter 7 — Obfuscation; Redaction; Data Masking; Privacy; Sensitive Information; Chapter 13 — Fitness for Purpose and Controlled Transformation.
===============
A report displaying birth date contains possible, but incorrect values. What is a possible explanation?
Birth date is populated from two source systems, both of which record the birth date in the birth date field
Birth date is populated from a single source system, which does not contain birth date
Birth date is populated from a single source system, which contains missing values
Birth date is populated from a single source system, where the date field is an offset value of 1601
Birth date is populated from two source systems, one of which stores marriage date in the birth date field
The critical wording is “possible, but incorrect values.†This describes values that satisfy basic syntactic or domain validation—they look like legitimate dates—but do not accurately represent the real-world attribute defined by the field.
If two systems contribute data and one maps marriage date into the birth-date field, the resulting values can be perfectly valid calendar dates while being semantically incorrect as birth dates. This is principally an Accuracy defect, because DAMA defines accuracy in terms of how correctly data represents the real-world object or event it is intended to describe. It may also expose a consistency and integration-mapping problem between source systems. DAMA's quality framework distinguishes accuracy from completeness: data can be populated and formally valid while still being factually wrong.
Missing values would primarily produce a Completeness defect rather than populated-but-incorrect values. Two correctly mapped systems would not inherently explain the problem. An offset or technical date representation could create transformation problems, but the scenario most directly illustrates semantic mis-mapping between data elements.
The appropriate remediation is therefore not simple cleansing alone. Metadata mappings, source-to-target specifications, lineage, business definitions, and integration rules should be corrected at the root cause.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Accuracy, Completeness and Consistency; Root-Cause Remediation; Data Profiling; Metadata Management; Data Integration and Interoperability.
===============
Database monitoring tools measure key database metrics, such as:
Create, read, normalization, user access
Capacity, availability, backup instances, data quality
Create, read, update, delete
Capacity, design, normalization, user access
Capacity, availability, cache performance, user statistics
DAMA-DMBOK2 identifies capacity, availability, cache performance, and user statistics as representative metrics captured by database-monitoring tools. Database monitoring is an operational control mechanism used by Database Administrators and platform teams to understand whether database infrastructure is available, adequately sized, responsive, and being used as expected. DMBOK2 explicitly describes monitoring tools as automating the observation of metrics such as capacity, availability, cache performance, and user statistics.
Capacity metrics indicate resource consumption and future growth requirements. Availability measures whether data services remain accessible when required. Cache-performance measures help identify inefficient access patterns and bottlenecks, while user statistics provide information about workload and database consumption.
The other choices mix design concepts, CRUD operations, or Data Quality concepts with operational monitoring measures. For example, normalization is a modelling technique rather than a routine runtime performance metric. Similarly, “create, read, update, delete†describes basic data operations rather than monitoring indicators.
Although database performance and Data Quality are distinct disciplines, poor operational performance can affect Data Quality dimensions such as timeliness and availability. Monitoring therefore supports the technical environment in which governed, reliable information is delivered to users and applications.
Reference Topics: DAMA-DMBOK2 — Data Storage and Operations; Database Operations; Capacity and Availability Management; Monitoring; Chapter 13 — Timeliness and Operational Fitness for Purpose.
===============
The search function associated with a document management store is failing to return known artefacts. This is due to a failure of:
Maintaining appropriate metadata on each document
Maintaining public access to all documents in the document management store
Effective data quality metrics
Data privacy and confidentiality procedures
Business intelligence implementation
A document-management search function depends heavily on metadata to identify, index, classify, and retrieve stored content. If known artefacts exist in the repository but cannot be found through search, inadequate or incorrectly maintained metadata is the most direct explanation.
DAMA-DMBOK2 distinguishes document content from the metadata that describes it. Typical descriptive metadata includes title, author, subject, keywords, document type, creation date, classification, and other characteristics used by retrieval mechanisms. Administrative and structural metadata may additionally support lifecycle management, versioning, access control, and relationships among document components. Without sufficiently accurate and complete metadata, the repository may physically contain a document while users remain unable to discover it.
This is also a Data Quality problem. Metadata itself is data and must satisfy quality expectations such as completeness, validity, consistency, and accuracy. A missing subject classification, incorrect document type, or inconsistent keyword convention can directly reduce findability.
Public access is neither required nor desirable for all documents, especially where confidential information exists. Business Intelligence is unrelated to basic document discovery, and Data Quality metrics alone do not make a document searchable unless the underlying metadata is correctly populated.
Reference Topics: DAMA-DMBOK2 — Document and Content Management; Metadata Management; Descriptive Metadata; Search and Retrieval; Chapter 13 — Metadata Quality and Fitness for Purpose.
===============
A customer table contains 50,000 records. Mandatory Tax Identification Number values are missing from 2,500 records. Which metric most directly measures the defect?
Consistency percentage
Completeness percentage
Uniqueness percentage
Reasonableness percentage
The defect is measured using Completeness. Completeness evaluates whether all required records or values that should be present are actually present. In this scenario, the mandatory Tax Identification Number is absent from 2,500 of 50,000 customer records.
A straightforward attribute-level completeness metric would be calculated as the number of populated required values divided by the number expected. Therefore, 47,500 of 50,000 records contain the required value, producing a completeness result of 95%.
Completeness does not establish that populated values are accurate. A record may contain a Tax Identification Number and therefore pass the completeness rule while containing the wrong number. This separation between completeness and accuracy is essential when designing Data Quality scorecards. DAMA-aligned guidance defines completeness as the presence of required values and explicitly distinguishes it from factual correctness.
Governance should also determine whether the attribute is genuinely mandatory for every customer type. Quality rules must reflect business applicability rather than blindly requiring every field for every record.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Completeness; Data Quality Metrics; Business Rules; Critical Data Elements; Scorecards.
===============
Which of the following is a reason why organisations do not dispose of non-value-adding information?
The organisation's data quality benchmark diminishes
The metadata repository can not be updated
The information is never out of date
Storage is cheap and easily expanded
Data modelling the content is hard to reproduce
The correct answer is storage is cheap and easily expanded. DAMA-DMBOK2 discusses retention and disposal as important lifecycle-management responsibilities. Although information that no longer provides business, legal, regulatory, historical, or evidentiary value should normally be disposed of according to approved retention policies, organizations often postpone disposal because modern storage appears inexpensive and technically easy to expand.
This reasoning is deceptive. The acquisition cost of storage is only one component of total information-management cost. Retaining unnecessary information also increases backup requirements, recovery time, discovery obligations, privacy exposure, security risk, metadata-management effort, migration complexity, and the volume of information that must be governed. DMBOK2 therefore stresses that non-value-adding information should not be retained merely because storage capacity is readily available.
From a Data Quality perspective, excessive retention also increases the population of obsolete, redundant, and potentially inconsistent information. This makes profiling, lineage analysis, master-data reconciliation, and authoritative-source identification more difficult.
A sound governance program consequently combines retention schedules, legal requirements, metadata classification, defensible disposal procedures, and clear accountability so that data is retained for legitimate reasons rather than technological convenience.
Reference Topics: DAMA-DMBOK2 — Document and Content Management; Retention and Disposal; Information Lifecycle; Data Governance; Chapter 13 — Data Quality and Obsolete Data.
===============
Multiple data warehouses, lakes, and swamps are an indicator that the organisation values:
Data modelling over data architecture
Data quality over master data management
Technology acquisition over a single point of truth
Processes over a single golden record
Security over enterprise data management
A proliferation of independently implemented data warehouses, lakes, and poorly governed “data swamps†commonly indicates that the organization has emphasized technology acquisition over establishing an integrated, authoritative view of enterprise data. The issue is not the existence of multiple technologies by itself; distributed platforms can be justified by legitimate architectural requirements. The warning sign is uncontrolled duplication without clear authoritative sources, enterprise architecture, governance, lineage, and consistency controls.
DAMA-DMBOK2 treats Data Architecture as the mechanism for aligning data assets with enterprise strategy and reducing unnecessary fragmentation. A mature architecture identifies authoritative sources, data flows, integration patterns, shared definitions, and the required relationship between operational and analytical stores. The DAMA framework also emphasizes that technologies should support business and data-management objectives rather than dictate them.
From a Data Quality perspective, uncontrolled proliferation creates competing versions of values, inconsistent transformations, duplicate master records, contradictory metrics, and unclear ownership. These conditions undermine consistency and make root-cause analysis difficult.
A “single point of truth†does not necessarily require one physical database. It requires controlled authority and consistent semantics across the landscape.
Reference Topics: DAMA-DMBOK2 — Data Architecture; Data as an Asset; Authoritative Sources; Data Integration; Metadata; Data Quality Consistency.
===============
Following the rollout of a data issue process, there have been no issues recorded in the first month. The reason for this might be:
The automatic deletion of all issues in the database
There are no data issues in the enterprise
Lack of credibility in the data governance process to effect changes
The denial of overtime requests
Staff staying back late to enter the issues into the system
A complete absence of reported issues immediately after introducing an enterprise Data Issue process is more likely to indicate lack of confidence in the governance process than genuinely flawless enterprise data. The certification material explicitly identifies lack of credibility in the process's ability to effect change as the plausible explanation.
Effective issue management depends on organizational trust. Employees need to believe that documenting an issue will lead to triage, ownership, escalation, root-cause investigation, remediation, and appropriate communication. If previous problems disappeared into a queue without action—or if raising defects creates organizational friction—users may simply stop reporting them.
DAMA-DMBOK2 treats Data Governance implementation as an organizational-change challenge rather than a purely procedural exercise. Governance must demonstrate authority, responsiveness, transparency, and measurable outcomes to establish credibility.
For Data Quality, issue volumes must also be interpreted carefully. “Zero issues reported†is not equivalent to “zero defects.†Complementary evidence should come from profiling, automated monitoring, quality metrics, user feedback, reconciliation, and operational outcomes.
Management should investigate reporting barriers, communicate resolved cases, establish clear escalation paths, and demonstrate that identified problems produce tangible improvements.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Governance Adoption and Credibility; Data Issue Management; Chapter 13 — Issue Identification, Escalation, Root-Cause Analysis and Remediation.
===============
Data Staging areas are populated from source databases using extract functions, transformed and then:
Made ready for review by a data modeller, for physical modelling
Made ready with advanced profiling and reported into a visualisation tool
Made ready for metrics calculations and key performance indicators
Made ready for metadata repository loading
Made ready and loaded into the data warehouse
The conventional flow is Extract, Transform, and Load (ETL). Data is extracted from source systems, placed in a staging environment where required transformations are performed, and then loaded into the target Data Warehouse. Consequently, the transformed staging data is made ready and loaded into the Data Warehouse. DAMA-oriented training material presents this exact progression.
The staging area provides a controlled environment in which source data can be integrated without imposing transformation workloads directly on operational systems. Typical activities include datatype conversion, standardization, code translation, matching, deduplication, validation, derivation, consolidation, and handling rejected records.
From a Data Quality perspective, staging is especially important because many quality controls can be applied before data reaches the analytical repository. Profiling can identify unexpected distributions, invalid values, missing records, duplicate entities, referential-integrity defects, and inconsistent reference values. Failed records can then be quarantined or remediated according to governed rules.
Although metrics and profiling may occur in staging, they are supporting activities rather than the final destination in the ETL sequence. Likewise, metadata repositories document the transformation rather than serving as the main target for the transformed business data.
Reference Topics: DAMA-DMBOK2 — Data Warehousing and Business Intelligence; ETL; Staging Areas; Data Integration; Chapter 13 — Profiling, Cleansing, Standardization and Validation.
===============
A minimal super key is:
Also known as a candidate key, it is a superkey without duplicated attributes.
Any set of attributes without duplicates that uniquely identifies an entity instance.
A synonym for a surrogate key.
A type of advanced index key structure, in the same family as Hash, Heap, B-Tree and Inverted.
Any set of attributes where each attribute that makes up the key is a foreign key in its own right.
An artificial key, made up of several meaningful components to help the reader understand the nature of the entity from the key alone.
A candidate key is a minimal super key. A super key is any set of attributes sufficient to uniquely identify an entity instance, but it may contain attributes that are unnecessary for uniqueness. A candidate key removes that redundancy: if any attribute is removed from the candidate key, the remaining attributes no longer uniquely identify the entity.
DAMA-DMBOK2 makes this distinction explicitly: a candidate key is a minimal set of one or more attributes identifying an entity instance, and “minimal†means that no subset of the candidate key can perform the same unique-identification function.
Option B is incomplete because a set can uniquely identify an entity while still containing unnecessary attributes; that would qualify as a super key but not necessarily a minimal super key. A surrogate key is different: it is an artificial identifier introduced primarily for technical identification. Foreign-key composition and physical index structures likewise do not define candidate-key minimality.
This concept has direct Data Quality implications. Correctly defined candidate and primary keys support uniqueness and integrity, prevent duplicate entity instances, enable reliable referential relationships, and improve entity matching within Master Data Management. Key definitions should also be captured as structural metadata so profiling and quality rules can consistently test duplicate and orphan conditions.
Reference Topics: DAMA-DMBOK2 Chapter 5 — Data Modeling and Design; Keys; Candidate Keys; Chapter 13 — Uniqueness and Integrity; Metadata Management; Master Data Management.
===============
The creation of overly complex enterprise integration over time is often a symptom of:
Multiple integration technologies
Multiple data warehouses
Multiple application coding languages
Multiple data owners
Multiple metadata tags
The strongest indicator is the uncontrolled proliferation of multiple integration technologies. Enterprise integration environments often become progressively more complex when different projects independently adopt different ETL platforms, messaging products, APIs, middleware technologies, replication mechanisms, file-transfer approaches, and proprietary interfaces.
The problem is architectural fragmentation. Each integration technology introduces its own configuration methods, transformation logic, monitoring mechanisms, metadata, failure handling, security controls, support skills, and operational procedures. As the number of technologies increases, point-to-point dependencies multiply and the organization accumulates technical debt. This makes end-to-end lineage, troubleshooting, change-impact analysis, and consistent application of Data Quality rules significantly harder. The source question itself identifies this specific integration-management issue. DAMA-oriented exam material likewise identifies multiple integration technologies as the characteristic cause of an overly complex integration landscape.
DAMA-DMBOK2 therefore emphasizes managed Data Integration and Interoperability architecture, reusable patterns, common models, governed interfaces, and metadata describing transformations.
From a Data Quality perspective, uncontrolled integration technology can result in inconsistent transformations, duplicated cleansing rules, mismatched reference values, and poorly understood lineage. Standardization reduces these risks by making movement and transformation processes more transparent and governable.
Reference Topics: DAMA-DMBOK2 Chapter 8 — Data Integration and Interoperability; Integration Architecture; Metadata and Lineage; Chapter 13 — Quality Controls Across Data Movement.
===============
An Orders table contains an order whose Customer_ID does not exist in the Customer table. Which Data Quality dimension is most directly violated?
Reasonableness
Completeness
Integrity
Currency
The primary issue is Data Integrity, specifically referential integrity. The Orders record references a Customer_ID that does not correspond to an existing Customer record, meaning the relationship defined by the data model has been violated.
Integrity concerns whether structural relationships and constraints among data elements remain valid. In relational environments, this commonly includes primary-key uniqueness, foreign-key relationships, mandatory relationships, and cardinality constraints.
A customer identifier could be syntactically valid and populated, yet still fail integrity because no corresponding parent entity exists. This illustrates why integrity is different from completeness or validity. The DMBOK2 dimension model explicitly associates integrity with unique identifiers, cardinality, and referential integrity concepts.
Remediation should determine why the orphan record occurred. Potential causes include incorrect load sequencing, deletion of the parent record, integration failure, transformation defects, or absence of database constraints.
Metadata and Data Modeling establish the expected relationship; Data Quality controls then measure whether operational data conforms to it.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Data Integrity; Referential Integrity; Chapter 5 — Relationships and Keys; Data Integration Controls.
===============
A Data Quality rule states that every active supplier must have a valid tax identifier. Who should approve whether the rule is mandatory for the Supplier domain?
Appropriate business governance/stewardship authority
Any developer working on the database
The network administrator
The storage vendor
Whether the rule is mandatory is a business governance decision and should be approved through the appropriate Data Owner, Data Steward, or governance authority responsible for the Supplier domain.
A developer can implement a database constraint, but technical implementation does not establish business policy. The organization must determine whether every active supplier genuinely requires a tax identifier, whether exceptions exist by jurisdiction or supplier type, and what consequences should follow when the rule is violated.
The approved rule should be documented as metadata and linked to the relevant business term, physical fields, authoritative sources, threshold, and responsible steward.
This illustrates the separation between governance and execution. Governance defines and authorizes the requirement; operational and technical teams implement the validation and monitoring.
DAMA's public framework places Governance across the Data Management knowledge areas and Metadata Management as the mechanism for maintaining definitions, lineage, and usage context.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Governance and Stewardship; Chapter 13 — Data Quality Rules; Metadata Management; Supplier Master Data.
===============
A Data Quality team repeatedly corrects invalid postal codes in the warehouse, but the same errors reappear after every nightly load. What is the most appropriate long-term response?
Increase the number of cleansing scripts in the warehouse
Correct the source or integration process causing the defect
Ignore the issue because cleansing already works
Reduce the frequency of data loads
The correct long-term response is to remove the root cause in the source or integration process. Repeated warehouse cleansing treats symptoms. If the defect is recreated every night, operational costs continue and downstream systems remain exposed until the cleansing process executes.
DAMA Data Quality practice emphasizes sustained improvement rather than repeated correction. The quality lifecycle therefore includes identifying defects, determining business impact, analyzing root causes, implementing corrective actions, and monitoring results. The current Chapter 13 revision specifically strengthens clarification of the Data Quality Improvement Lifecycle and responsibilities within it.
The underlying cause might be weak source validation, an incorrect source-to-target mapping, outdated Reference Data, transformation logic, or missing governance over postal-code rules.
Cleansing remains appropriate where historical data must be repaired or immediate downstream protection is required. However, preventive control should be introduced as close to the creation point as practical.
Metadata and lineage help trace the postal code through the data flow, while governance establishes who has authority to change the offending process.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Root-Cause Analysis; Remediation; Prevention; Cleansing; Data Integration and Lineage.
===============
A Data Quality incident affects several downstream systems. What should be determined before individual teams independently correct their copies?
The root cause and authoritative point of remediation
Which team has the largest budget
Which system contains the most records
Which copy is easiest to delete
The organization should determine the root cause and authoritative remediation point before multiple teams independently modify downstream copies.
If a defective value originates in a source system and is distributed to five downstream applications, correcting those five copies independently may provide temporary relief but leaves the upstream defect intact. The incorrect value may simply be redistributed again.
Lineage should be used to trace the defect through source systems, transformations, integration layers, Master Data hubs, warehouses, and reports. Once the origin is understood, governance can determine which system or process is authoritative and where correction should occur.
This approach minimizes inconsistent local fixes and reduces the risk that different teams apply incompatible interpretations of the same issue.
DAMA's Chapter 13 revision explicitly strengthens the relationship between Data Quality and Metadata Management, Reference/Master Data Management, Modeling, and Data Integration. Those connections are precisely what enable cross-system issue resolution.
Downstream correction may still be required after the authoritative source is fixed, but it should be coordinated.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Root-Cause Analysis; Issue Management; Metadata Lineage; Authoritative Sources; Data Integration.
===============
A Data Quality team finds that 78% of customer complaints are caused by four defect categories out of twenty recorded categories. Which tool is most appropriate for prioritizing improvement activity?
Pareto analysis
Encryption analysis
Normalization analysis
Network segmentation
Pareto analysis is the appropriate technique because it identifies the relatively small number of causes responsible for a large proportion of observed problems.
The team should arrange defect categories in descending order by frequency or business impact and calculate their cumulative contribution. If four categories account for 78% of complaints, prioritizing those categories is likely to generate substantially greater benefit than spreading equal effort across all twenty.
The value of Pareto analysis in Data Quality is not that every problem follows an exact 80/20 relationship. Rather, it provides an evidence-based mechanism for focusing scarce remediation resources on the defect classes that generate the most significant outcomes.
Frequency should not be the only prioritization criterion. A rare defect could produce severe regulatory, safety, financial, or reputational consequences. Governance should therefore combine volume analysis with business impact and risk.
Once priority defects are selected, the team should perform root-cause analysis and introduce preventive controls instead of merely fixing individual records.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Statistical Quality Tools; Pareto Analysis; Issue Prioritization; Business Impact; Root-Cause Remediation.
===============
When a data quality team has more issues than they can manage, they should look to:
Establish a program of quick wins targeting easy fixes over a short time period
Delete any issue that is greater than 6 months old
Establish a program of quick wins targeting easy fixes over a short time period
Initiate data quality improvement cycles, focusing on achieving incremental improvements
Hire more people
When the Data Quality issue backlog exceeds the team's practical capacity, the appropriate approach is to establish a program of quick wins targeting manageable fixes over a short period. The objective is to reduce immediate pressure, demonstrate measurable value, and create organizational momentum rather than attempting to attack every known defect simultaneously.
DAMA-DMBOK2 recommends a hybrid implementation approach: top-down support provides sponsorship, consistency, and resources, while bottom-up work identifies what is actually broken and produces incremental successes. It further emphasizes prioritization of remediation based on business impact, documented costs and benefits, and sustained issue-management operations.
Deleting old issues is arbitrary and can conceal unresolved business risk. Hiring additional personnel may increase capacity but does not correct weak prioritization. Data-entry validation is valuable preventive control, but it addresses only certain causes of defects.
Option D sounds plausible because continuous improvement is fundamental to Data Quality Management; however, the immediate problem in this scenario is an unmanageable backlog. Short, prioritized wins provide the tactical mechanism for regaining control while the broader continuous-improvement lifecycle continues.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Data Quality Implementation Guidelines; Prioritization; Remediation; Incremental Successes; Root-Cause Analysis; Continuous Improvement.
===============
Who should normally define the business meaning and acceptable quality requirements for a critical customer attribute?
A Business Data Steward working with relevant stakeholders
The network administrator alone
The database optimizer alone
The backup operator alone
A Business Data Steward, working with relevant business stakeholders and governance bodies, is the appropriate role to define or coordinate business meaning and quality expectations.
Quality requirements cannot be derived solely from database structures. The organization must understand what the attribute means, how it is used, what values are acceptable, which source is authoritative, and what consequences arise when it is wrong.
Data Stewards bridge business knowledge and formal Data Management controls. They commonly participate in maintaining glossary definitions, defining business rules, resolving issues, clarifying ownership, and establishing Data Quality expectations.
Technical specialists contribute implementation knowledge. A DBA may identify datatype constraints, while an integration specialist can implement transformations. Neither should independently determine business meaning.
DAMA's public framework describes Metadata Management as supporting definitions, lineage, and governance, while Reference and Master Data Management ensures consistency in shared core entities.
The strongest operating model therefore combines business accountability with technical execution.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Data Stewardship; Chapter 11 — Business Metadata; Chapter 13 — Data Quality Requirements; Critical Data Elements.
===============
Discovering and documenting metadata about physical data assets provides:
An estimation of balance sheet value of enterprise data
Scoping boundaries of the data dictionary
Insights into the temporal data quality
Information on how data is transformed as it moves between systems
Effective project scope management
Discovering and documenting metadata about physical data assets provides visibility into how data moves and is transformed between systems. This is a core function of technical metadata and data lineage.
DAMA-DMBOK2 states that technical metadata describes the technical characteristics of data, the systems that store it, and the processes that move it within and between systems. Examples include physical table and column names, ETL job information, source-to-target mappings, and lineage documentation with upstream and downstream impact information. The certification question therefore points directly to option D; the same interpretation is independently reflected in published CDMP-oriented material.
This capability is critical to Data Quality because a defect observed in a report or downstream application may not originate in the current system. It may have been introduced through extraction logic, transformation rules, reference-data conversion, aggregation, or loading.
Documented physical metadata allows analysts to trace the affected value backward, identify its source, inspect each transformation, and isolate the point at which the defect occurred. It also supports change-impact analysis: modifying a source field or mapping can reveal which downstream assets may be affected.
Metadata therefore provides the evidence required for systematic root-cause analysis rather than symptom-based correction.
Reference Topics: DAMA-DMBOK2 Metadata Management — Technical Metadata; Source-to-Target Mapping; Data Lineage; Data Integration and Interoperability; Chapter 13 — Root-Cause Analysis and Quality Monitoring.
===============
A source application is modified so that a mandatory product identifier can no longer be left blank. Which type of Data Quality action is this?
Preventive control
Detective control only
Historical archiving
Data obfuscation
The application change is a preventive Data Quality control because it prevents a known defect from being created in the first place.
A mandatory-field control ensures that records cannot be accepted without the required Product Identifier. This differs from a detective control, which would identify missing identifiers after records had already entered the system.
Preventive controls are generally preferable where the organization controls the point of data creation and the business rule is sufficiently clear. They reduce downstream cleansing, exception handling, reconciliation, and operational rework.
However, making the field technically mandatory is appropriate only if the business requirement genuinely applies to every relevant record. Governance and stewardship should confirm the rule before implementation. If certain product categories legitimately lack the identifier, the validation should incorporate those conditions instead of enforcing an overly broad requirement.
The Product Identifier may also link the transaction to governed Product Master Data, making integrity and consistency important alongside completeness.
DAMA's Data Quality framework emphasizes defining rules, detecting defects, implementing improvement, and integrating quality with Governance and Master Data disciplines rather than relying solely on after-the-fact cleansing.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Preventive Controls; Completeness; Data Quality Improvement Lifecycle; Chapter 10 — Product Master Data; Governance.
===============
A business case for adding a new master data management solution is dependent on achieving greater value from:
Mining golden records
Decentralising shared data as it is closer to where the business needs it
Centralised coordination of shared data vs the data management and integration cost of running another copy of the data.
Avoiding reference data management
Decommissioning all the other systems that manage this data
An MDM investment is justified when the business value obtained from centralized coordination of shared data exceeds the additional cost of managing and integrating another controlled representation of that data. This is the economic trade-off expressed by option C and is also the intended answer in published versions of this DAMA question.
DAMA-DMBOK2 positions Master Data Management as a mechanism for improving consistency and control over enterprise entities such as customers, products, suppliers, employees, and locations. Uncoordinated local copies create redundant maintenance, inconsistent definitions, duplicate records, reconciliation costs, and integration complexity. DAMA sources emphasize that controlling shared master and reference data reduces both cost and risk generated by inconsistencies between systems.
However, introducing an MDM repository also incurs costs: integration, stewardship, matching, survivorship logic, synchronization, governance, maintenance, and potentially another physical copy of shared information. The business case must therefore demonstrate that improved coordination, quality, reuse, reporting, and operational consistency outweigh those costs.
MDM does not require all contributing systems to be decommissioned, nor does it eliminate Reference Data Management. Similarly, “mining golden records†is not the fundamental economic basis for MDM.
Reference Topics: DAMA-DMBOK2 Chapter 10 — Master Data Management; Business Drivers; Shared Data; Central Coordination; Data Integration; Data Quality and Stewardship.
===============
What is the purpose of Data Governance?
Ensure that financial performance of the company is improved
Encompass the entire lifecycle of a data asset
Ensure an organization gets value out of its data
Establish processes and functions through which data can be enabled for use and also maintained
To ensure that data is managed properly, according to policies and best practices
The precise DAMA-DMBOK2 purpose of Data Governance is to ensure that data is managed properly according to policies and best practices. DMBOK2 distinguishes this governance purpose from the broader objective of Data Management, which is concerned with obtaining value from data throughout its lifecycle.
Governance establishes the authority and control framework within which operational data-management functions work. This includes strategy, policy, standards, accountability, stewardship, compliance, issue escalation, quality oversight, and decision rights. DMBOK2 specifically identifies governance responsibilities involving policies for metadata, access, security and quality; setting Data Quality and Data Architecture standards; and providing oversight and corrective action.
This distinction is critical for Data Quality. Data Quality Management performs activities such as profiling, measurement, root-cause analysis, cleansing and monitoring. Data Governance determines who has authority, which requirements are mandatory, what thresholds are acceptable, which issues require escalation, and who owns remediation decisions.
Metadata Management supports governance by recording authoritative definitions, lineage, ownership and quality rules. Master Data Management operationalizes governance for shared business entities and reference domains.
Options B, C and D describe legitimate aspects of wider Data Management, but they do not state the specific purpose assigned to Data Governance by DMBOK2.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Introduction; Goals and Principles; Governance Oversight; Chapter 13 — Data Quality and Governance.
===============
Critical Data is most often used in:
Regulatory, financial, or management reporting
Business operational needs
Measuring product quality and customer satisfaction
Business strategy, especially efforts at competitive differentiation
All of these
All of these is correct because DAMA-DMBOK2 defines criticality according to the business use and impact of the data rather than by a particular technical characteristic. Critical Data Elements are commonly associated with regulatory, financial, or management reporting; operational business needs; measurement of product quality and customer satisfaction; and strategic initiatives such as competitive differentiation. DAMA's revised Chapter 13 also explicitly adds and clarifies the concept of Critical Data Elements within the Data Quality framework.
The purpose of identifying critical data is prioritization. Organizations cannot apply identical levels of profiling, monitoring, stewardship, remediation, and control to every data element. Data whose failure could create substantial operational, financial, legal, regulatory, customer, or reputational impact receives greater management attention.
For example, values feeding regulatory reports require stringent accuracy and traceability; operational data may require high availability and timeliness; customer-satisfaction measures require reliable and consistent inputs; and strategically important data can affect competitive decisions. Current DAMA-aligned material confirms these same usage categories for Critical Data Elements.
Once Critical Data Elements are identified, they should be linked to business definitions, Data Owners and Stewards, quality dimensions, measurable rules, lineage, authoritative sources, thresholds, and issue-management processes.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Critical Data Elements; Business Drivers; Data Quality Requirements; Prioritization; Data Governance; Metadata Management.
===============
What is Data Stewardship?
A prioritized program of work with scoped boundaries
The creation of compelling vision for Data Management across the enterprise
Refers to the role responsible for creating policies, procedures, and rules that govern data in the organization
A collection of tools that ensure an organization's privacy policy
A position accountable and responsible for data and processes that ensure effective control and use of data assets
DAMA-DMBOK2 defines Data Stewardship in terms of accountability and responsibility for data and for the processes that ensure its effective control and use. The wording in option E closely matches the DMBOK2 definition: stewardship describes accountability and responsibility for data and processes that ensure the effective control and use of data assets.
Stewardship may be formalized through named positions or responsibilities embedded within existing business roles. Its practical scope commonly includes managing business terminology, defining valid values and business rules, establishing or approving Data Quality requirements, resolving data issues, applying standards, and supporting governance decisions.
Option C is too narrow and also confuses stewardship with the broader policy-setting responsibilities of Data Governance. Stewards participate in developing and implementing policies and standards, but stewardship is not merely a policy-creation role. Similarly, it is not a privacy technology function or a project-management construct.
Within Data Quality Management, Data Stewards are critical because quality must be defined relative to business requirements. They help determine what “fit for purpose†means, establish acceptable thresholds, prioritize defects according to business impact, and participate in root-cause remediation.
Metadata Management records stewardship decisions, while Master Data Management uses those decisions to govern shared enterprise entities.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Data Stewardship; Steward Responsibilities; Business Glossary; Chapter 13 — Data Quality Governance and Issue Management.
===============
Reference data is often a list of code values with their full names. One example is:
Person age associated with accessibility requirements
State or province associated with the related country
Encrypted code values with the unencrypted data value
Country codes associated with the country names
Master data values encoded into reference values
Country codes associated with country names are a classic example of Reference Data. DAMA-DMBOK2 defines Reference Data as data used to characterize or classify other data or to relate organizational data to externally defined information. The simplest reference-data structure consists of a code and its corresponding description. DMBOK2 specifically uses geographic and standards-based examples, including country codes such as DE, US, and TR.
A code such as GB therefore acts as a standardized machine-processable value, while “United Kingdom†supplies the human-readable meaning. Such code lists are reused across applications, integrations, reporting platforms, Master Data systems, and analytical environments.
Reference Data quality is important because inconsistent code sets can create widespread downstream defects. If different systems use incompatible country codes, the organization may experience failed integrations, incorrect aggregation, duplicate mappings, reporting discrepancies, and inconsistent interpretation. Data Governance should therefore establish authoritative sources and stewardship responsibilities for important reference domains.
Metadata Management records the meaning, provenance, permitted values, relationships, and mappings between code sets. Master Data Management frequently consumes these governed codes to classify master entities such as customers, suppliers, products, and locations.
The other choices describe relationships or transformations, but they do not represent the straightforward code-value/description list structure identified by DAMA.
Reference Topics: DAMA-DMBOK2 Chapter 10 — Reference and Master Data; Reference Lists; Code Sets; Authoritative Sources; Chapter 13 — Validity, Consistency and Integrity.
===============
A Data Quality process has remained within statistical control limits for six months, but management wants the average defect rate reduced further. What is the most appropriate conclusion?
A stable process may still require deliberate process improvement
Statistical stability proves the process is perfect
All monitoring should stop
Every record outside the average must be deleted
A process can be statistically stable yet still perform at an unacceptable quality level. Statistical control means that variation is predictable within the established process; it does not mean that the average level of defects satisfies business expectations.
For example, a process may consistently produce a 3% defect rate with very little month-to-month variation. If the business requirement is below 0.5%, the process is stable but incapable of meeting the desired quality level without improvement.
The appropriate response is therefore deliberate process improvement aimed at changing the underlying process and reducing its baseline defect rate. This differs from reacting to random individual observations within normal control limits.
After improvement, new performance data should be collected and control parameters reevaluated once the revised process reaches a stable state.
DAMA's revised Chapter 13 explicitly maps Shewhart and Deming improvement-cycle stages to Data Quality processes and preserves Statistical Process Control as supporting material.
The central principle is that control and capability are different questions: control asks whether the process is stable; business quality requirements determine whether that stable performance is good enough.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Statistical Process Control; Process Stability; Continuous Improvement; Shewhart/Deming Cycle; Thresholds.
===============
Two departments use the term "Active Customer" but apply different definitions. Which Data Management capability should be addressed first?
Business glossary governance
Storage compression
Index rebuilding
Network segmentation
The immediate requirement is Business Glossary governance. The issue is semantic: two departments use the same term but assign different meanings to it.
A governed business glossary establishes agreed terminology, definitions, synonyms, related concepts, ownership, and usage context. Resolving the definition does not necessarily mean forcing every business process to use one operational rule; there may be legitimate variants. The critical requirement is that those variants be explicitly named and documented so consumers understand the distinction.
Metadata Management provides the structures needed to preserve and distribute these definitions. DAMA's public Metadata Management guidance identifies business glossaries and data dictionaries as core mechanisms for improving shared understanding.
This semantic clarity is fundamental to Data Quality. A completeness or accuracy score for "Active Customer" is meaningless if departments measure different populations under the same label.
After the definition is governed, technical metadata should link the glossary concept to physical fields, calculations, reports, and quality rules.
Reference Topics: DAMA-DMBOK2 — Business Glossary; Metadata Management; Data Stewardship; Semantic Consistency; Data Quality Rules.
===============
TESTED 11 Oct 2026