Glossary

What is data deduplication?

Matches, merges and cleans scattered records into one trusted golden master, with human stewardship and full audit trails.

What is data deduplication?

Data deduplication is the systematic process of identifying, matching, and merging duplicate records scattered across multiple source systems into a single, authoritative version known as a golden master. Rather than allowing conflicting or redundant entries to accumulate and erode data quality over time, deduplication consolidates them into one trusted record while preserving the complete history of where each piece of information originated. The result is a dataset that is clean, consistent, and reliable enough to serve as the foundation for reporting, procurement, finance, and supply chain operations — without the ambiguity that fragmented records introduce. Allmaz approaches deduplication through a deterministic matching engine that applies rule-based logic to identify duplicate candidates with precision and transparency. Critically, no AI-generated proposal is ever published without explicit human approval — a data steward reviews and confirms every merge decision before it takes effect. All edits are stored as overrides rather than applied in place, meaning the original data is never destroyed and every change remains fully reversible. Combined with per-record source lineage and a complete audit trail, this architecture gives organisations the confidence to consolidate their data at scale while remaining in full control of every decision made along the way.

Capabilities

Why data deduplication matters

Eliminates redundant records so every team works from a single, reliable source of truth rather than fragmented, conflicting data spread across disconnected systems.

Reduces downstream errors in reporting, procurement, and operations that stem directly from duplicate or mismatched entries reaching business processes unchecked.

Saves significant time previously spent on manual reconciliation by automating duplicate detection and candidate grouping through a deterministic matching engine.

Maintains full source lineage per record so stakeholders always know exactly where data originated, which sources contributed to a merge, and how the golden master was constructed.

Supports compliance and governance requirements through a complete, timestamped audit trail that logs every merge, override, and approval decision against a named user.

Preserves operational flexibility with a reversible-by-design approach — no original data is ever overwritten, and any edit or merge can be undone at any point without data loss.

Key features of Allmaz data deduplication

Deterministic matching engine

Records are compared using rule-based, deterministic logic that identifies duplicates with precision, reducing false positives and giving human stewards clear, explainable match candidates to review.

Golden master record

Matched duplicates are merged into one authoritative golden master that represents the best, most complete version of each entity — eliminating ambiguity across connected systems.

Human stewardship and approval

No AI-generated proposal is published without explicit human approval. Data stewards remain in control of every merge decision, ensuring accountability at every step.

UNSPSC taxonomy classification

Records are automatically classified into the internationally recognised UNSPSC taxonomy, enabling consistent categorisation across procurement, finance, and supply chain workflows.

Reversible edits and override storage

All changes are stored as overrides rather than in-place modifications, meaning the original data is never destroyed and any edit can be reversed at any time.

Full audit trail and source lineage

Every record carries a complete audit trail showing who changed what and when, alongside source lineage that traces each data point back to its origin system.

How data deduplication works in Allmaz

1Scattered records are ingested from multiple source systems and normalised into a consistent format ready for comparison.
2The deterministic matching engine analyses the normalised records and groups likely duplicates into candidate clusters based on defined matching rules.
3Each candidate cluster is presented to a human data steward for review, who confirms, adjusts, or rejects the proposed merge.
4Approved merges are consolidated into a single golden master record, with all contributing sources documented in the lineage log.
5Records are classified into the UNSPSC taxonomy to ensure consistent categorisation across business functions.
6All changes are saved as overrides in the audit trail, keeping the original data intact and every decision fully traceable.

Frequently asked questions about data deduplication

What is a golden master record and how is it created?

A golden master is the single, authoritative version of a record produced by merging duplicate entries identified across multiple source systems. It represents the most complete and trusted view of that entity within your dataset. The golden master is only created after a human data steward has reviewed the proposed merge and given explicit approval — it is never generated automatically.

Can merged records be undone if a mistake is made?

Yes. Allmaz is reversible by design. All edits are stored as overrides rather than applied directly to the original data, so any merge or change can be reversed at any time without data loss. The underlying source records remain intact throughout the entire process.

What role does a human steward play in the deduplication process?

Human stewards review every match proposal generated by the deterministic engine before it is published. No change is applied automatically — steward approval is required at each decision point. This ensures that a qualified person is accountable for every merge, adjustment, or rejection, and that no data is altered without deliberate human oversight.

What is UNSPSC classification and why is it part of deduplication?

UNSPSC (United Nations Standard Products and Services Code) is an internationally recognised taxonomy for categorising products and services. Incorporating UNSPSC classification into the deduplication workflow ensures that once records are consolidated into a golden master, they are also consistently categorised — enabling comparable, structured data across procurement, finance, and supply chain functions from the moment a record is created.

How does the audit trail support governance and compliance requirements?

Every action taken on a record — including merges, overrides, and steward approvals — is logged with a timestamp and user reference, creating an unbroken chain of accountability. Source lineage is also maintained per record, tracing each data point back to its origin system. Together, these capabilities provide the transparency and traceability needed for internal governance reviews, external audits, and regulatory compliance obligations.

Ready to build a trusted data foundation?

Explore how Allmaz can help your team match, merge, and govern records with confidence — combining deterministic precision with human oversight and a complete audit trail.

Request a demo