Algoscale

Case Study · Telecommunications

Employee MDM & Data Governance Design

Six-week advisory engagement designing HR golden record architecture, data governance, and DPDPA-aligned MDM evaluation for an Indian telecom group.

Client: Indian Telecom Group — Global Conglomerate Portfolio

19
design artefacts delivered across three workstreams in six weeks
125 attributes
canonical employee dictionary across 10 domains — each classified, owned, and quality-gated
9 systems
confirmed source estate after reconciliation — down from an initial count of 13
6 personas
role-based view specifications in the Employee Rich Profile serving layer

Source systems unified

Enterprise HCM Platform Central HR system of record
Payroll System Compensation and deductions
Attendance Management System Time, leave, and overtime tracking
Learning Management System Training and certification records
Background Verification Platform Compliance screening records
Recognition & Engagement Systems Employee sentiment and rewards data

The problem: nine systems, no single source of truth

The client is the Indian telecommunications arm of a global conglomerate whose portfolio spans several Fortune 500 companies. The telecom group operates three legal entities under a single HR function, with a workforce of under ten thousand active employees plus historical records going back several years.

Nine live systems covered the employee lifecycle: a central enterprise HCM platform, plus separate systems for payroll, attendance, learning, recognition, engagement surveys, background verification, and overtime tracking. Each system worked. None of them could reliably answer a basic question about a specific employee: what is true about this person right now?

Different systems held different versions of the same attribute. Ownership of any given field was assumed rather than assigned. Department codes diverged across platforms. Designation records went stale. When two systems agreed on an employee’s location and a third disagreed, there was no documented rule for which one to trust.

Two external pressures arrived at the same time. India’s Digital Personal Data Protection Act was moving from statute to enforcement, and the group needed to demonstrate attribute-level control over personal data it had never formally classified. And the group’s broader data strategy was pushing toward an Employee Rich Profile — a unified serving layer that HR business partners, managers, and the leadership team could query confidently. Neither objective was achievable without first knowing what the estate actually contained and who was responsible for it.

Algoscale was engaged to design the foundation — not to build it. The brief was explicitly advisory: produce the blueprints, the governance structures, and the platform recommendation that would allow a Phase 2 build to proceed without re-discovering the problem.

Three workstreams, six weeks, nineteen artefacts

The engagement ran three parallel workstreams, each producing its own set of deliverables, all anchored to a single canonical attribute dictionary.

Employee Rich Profile. A read-only serving layer over the golden record — designed for consumption, not administration. Deliverables: a canonical data model for the profile, role-based view specifications for six personas (ranging from the employee themselves through HR business partners to the leadership team), field-level attribute mapping to source systems, dashboard wireframes, integration architecture, and privacy and consent control design.

HR Data Governance Framework. The ownership and operating model for the attribute estate. Deliverables: a current-state system and process assessment, a data domain map and dictionary, an attribute sensitivity classification schema aligned to DPDPA, data quality rules and scoring standards, a per-function DQ check template, a governance policy suite, a RACI framework, and a governance council charter.

Golden Record and MDM. The technical and platform architecture for mastering the employee record. Deliverables: source system inventory and profiling, matching and survivorship design, golden record design and reconciliation approach, data lineage and audit trail architecture, a scored MDM platform evaluation, and a phased implementation roadmap.

Nineteen artefacts. Six weeks. The scope was achievable precisely because everything was anchored to one attribute list from day one — the data governance consulting discipline of establishing a single vocabulary before producing any downstream artefact.

The approach

Reconcile the estate before modelling it. The first substantive output was not a data model — it was an accurate count of systems. This sounds trivial. It was not. See finding two below.

Anchor everything to one attribute dictionary. A single canonical dictionary of 125 attributes across 10 domains became the spine of the engagement. Sensitivity classification, data quality rules, RACI assignments, consent design, and golden record survivorship logic all reference the same attribute list. Nothing was permitted to define its own parallel vocabulary. Parallel vocabularies are how governance frameworks quietly stop being used. The argument for why Customer 360 (or Employee 360) projects fail for exactly this reason is made in detail in Customer 360 Without MDM Is Cosplay.

Treat DPDPA as a design input, not a compliance appendix. India’s Digital Personal Data Protection Act was not relegated to a separate compliance section. Consent requirements, purpose limitation, and nomination rights were mapped to specific attributes and specific systems — so the schema reflected the law as written, not a GDPR template adapted to fit.

Be explicit about what is designed versus what is built. Every artefact carried this distinction on its face. Design completeness is real value. Presenting it as operational capability is not. Credibility in a Phase 2 build depends entirely on Phase 1 precision about what has and has not been delivered.

Five findings worth generalising

1. Spreadsheets were the de facto system of record for 14 attributes

Fourteen canonical attributes had no system of record other than a spreadsheet maintained by a named individual. These were not fringe fields — several fed compliance reporting processes.

This is the single most reusable finding from the engagement. Every large HR estate has these attributes. They are invisible in an architecture diagram, they survive system migrations untouched, and they are the reason “we already have an HCM platform” is never a complete answer to the master data question. Any MDM business case should start by counting them. The count is the business case.

2. A system count is a claim about platforms, not about labels

The estate initially appeared to contain thirteen systems. After verification it contained nine. Several apparent systems turned out to be different internal labels for modules of the same underlying platform. Two reporting tools had been counted as sources rather than consumers. One system was sunsetting. One was mid-migration. One was a downstream portal with no system-of-record status.

The correction mattered enormously. An inflated source count distorts integration effort, reconciliation scope, and licensing assumptions simultaneously. Every downstream number — interface count, survivorship complexity, MDM runtime estimates — had to be calculated against the correct nine, not the assumed thirteen.

The operating rule we now apply from this engagement: no system name or count enters a client-facing document until it has been confirmed in writing by the platform owner. Verbal confirmation in a workshop is not sufficient, because participants describe their own interface, not the platform behind it.

3. DPDPA has no equivalent to GDPR’s special categories

The sensitivity classification schema was initially drafted with a “special category” column carried over from GDPR practice. India’s DPDPA contains no statutory special-category framework — it treats personal data as a single class, with obligations attached to the processing purpose and consent mechanism rather than to the data type.

The column was removed. More importantly, the reason it was absent was documented as a deliberate, affirmative reading of the applicable law — not as an omission. An auditor encountering a classification schema with no special-category treatment will otherwise read it as an oversight. The documentation is the deliverable.

4. Completeness should be down-weighted in data quality scoring

Standard data quality frameworks weight completeness, accuracy, consistency, timeliness, and validity roughly equally. The framework produced for this engagement weighted completeness materially lower.

The reasoning: a missing field is a caught error. It is visible, it fails a null check, and someone will eventually chase it. A field that is populated but wrong — a stale designation, an unupdated location, a duplicated verification status — is a silent error. It passes every completeness test and propagates downstream into decisions undetected.

Scoring these error types equally systematically overstates dataset health. The hybrid-weighted model developed for this engagement pushes the quality score toward the errors that actually cause harm in downstream HR processes and regulatory submissions. The data management services practice has applied this weighting approach across subsequent engagements.

The weighting model was embedded in a single Excel-based DQ check template, parameterised with the specific attributes belonging to each HR domain. A payroll owner running the template sees only payroll-domain attributes and their associated quality rules; an L&D owner sees learning and certification fields; an HR operations owner sees core employment and designation fields. The attribute scope per function is pre-populated from the canonical dictionary, so the template is always in sync with the agreed attribute list.

The checks run on-demand rather than on a schedule. Different functions have different quality rhythms — payroll runs checks before the monthly close, L&D before annual review cycles, HR ops before a regulatory submission. A scheduled central batch check creates no accountability for the function that should have caught the error before it propagated. An on-demand template in the hands of the domain owner creates that accountability.

The Excel format was a deliberate Phase 1 choice. No new tooling to procure, no IT dependency for a function head who wants a spot check ahead of a deadline. The check logic — rules, attribute scope, and scoring formulae — is now defined, tested, and documented. A Phase 2 DQ platform replaces the transport layer, not the rules. That is the correct sequencing. Embedding the rules in a platform before you have chosen it guarantees a migration the moment you do.

5. Classification tier is constrained by handling reality

An attribute cannot be classified at a strict need-to-know tier if the organisation’s own practice handles it openly. Blood group is the clean example from this engagement: genuinely sensitive in principle, and displayed on ID cards, emergency rosters, and desk labels in practice.

Classifying it as Restricted would have produced a schema that was immediately and universally violated, teaching everyone that the classification scheme is decorative. It was classified one tier lower, at Confidential, with the handling reality documented alongside the rationale.

A classification scheme that reflects actual practice can be tightened incrementally. One that is ignored from day one cannot be recovered. The governance design philosophy applied throughout this engagement: design for the organisation that exists, document the gap, create a path to close it.

Platform recommendation

The MDM evaluation scored native and independent platforms against the client’s specific attribute set, integration estate, and stewardship maturity. Criteria included integration complexity, time to first golden record, stewardship tooling, total cost of ownership, vendor lock-in profile, and extensibility beyond the HR domain.

The recommendation was the enterprise HCM vendor’s own data management product as the primary path — lowest integration friction, existing commercial relationship, no new vendor onboarding or security review. Two independent platforms were named as best-value alternatives should the group prefer a vendor-neutral master data layer that could later extend beyond HR to cover other entity types across the conglomerate’s Indian entities.

The evaluation was scored against explicit criteria rather than expressed as a preference, so the recommendation could be defended to procurement and technology governance committees without reference to the authors.

What the engagement delivered

The group entered Phase 2 planning with the ambiguity removed.

A confirmed nine-system source estate replaced the assumed thirteen. Every integration estimate, reconciliation design decision, and MDM runtime assumption is now calculated against verified ground truth.

A 125-attribute canonical dictionary with DPDPA-aligned sensitivity classification, domain ownership, and data quality rules is the single vocabulary for every downstream workstream. HR business partners, data engineers, the privacy team, and the governance council all reference the same list.

Named owners for every attribute domain — assigned through the RACI framework and formalised through the governance council charter — means field-level accountability exists before the MDM build starts. Most implementations that fail in Phase 2 fail because ownership was assumed in Phase 1.

A scored MDM platform recommendation with a phased implementation roadmap gives the technology team a decision they can execute and a sequence they can resource, rather than a shortlist and an instruction to evaluate.

Golden record matching and survivorship design defines the rules for resolving conflicts between systems before any pipeline is written. The logic is documented, reviewed, and agreed; it does not have to be reverse-engineered from a model that has already run.

The argument that underpins this engagement — that master data is a governance problem before it is a technology problem — is made directly in Customer 360 Without MDM Is Cosplay. The employee record is the HR version of the same failure mode: the platform exists, the integrations run, and the data is still wrong because ownership and survivorship logic were never defined.

For the implementation counterpart — what a leaner MDM pattern looks like when the organisation is ready to build — see Serverless MDM: Lambda + Postgres on AWS. The governance work done in a design engagement like this one is what makes that build tractable rather than another shadow-schema migration.

Our data governance consulting practice runs advisory engagements of this kind — attribute dictionaries, sensitivity classification, ownership frameworks, platform evaluation — for organisations that need the blueprints before they can responsibly commission the build.

The stack

  • 125-attribute canonical dictionary (10 domains)
  • DPDPA-aligned sensitivity classification schema
  • Hybrid-weighted data quality rules
  • Golden record matching & survivorship design
  • Employee Rich Profile serving layer
  • 6-persona role-based view specifications
  • Governance council charter & RACI framework
  • MDM platform evaluation (scored)
  • Phased implementation roadmap

Related work

Want the same for your estate?

45 minutes, your source list, your reporting pain. We'll map the warehouse shape and show you what S.C.A.L.E. ships against it.

Book a walkthrough

Pick your starting point

Two quick diagnostics for the two questions we get most

No sales calls required to get real answers. Both tools return dedicated output in under 5 minutes.