The Rp 118,000 Compound
A customer checks out once through Tokopedia with an email address, then returns weeks later through a WhatsApp click-to-chat link that captures only a phone number. Because the CRM, the POS loyalty database, and the loan or order system each mint their own row — and none of the three enforces the NIK as a primary key — one human becomes three billing entities overnight. Call this a key vacuum, not carelessness: no amount of staff diligence repairs an intake architecture that never asks for the one identifier the state already issued. Every new channel widens the vacuum, which is why the tax recurs instead of fading.
The first layer is wasted messaging. Each duplicate generates 6-14 redundant outbound touches per year, and at Telkomsel's enterprise SMS rate of roughly Rp 45 per message alongside WhatsApp Business API utility-template rates near Rp 350, messaging waste alone runs Rp 900-Rp 4,300 per duplicate annually — the width of that band is pure channel mix, SMS-heavy programs sitting at the floor and WhatsApp-heavy ones near the ceiling. The second layer is human reconciliation: a fully loaded Jakarta contact-center agent costs about Rp 42,000 per hour, resolving one conflicting-record ticket takes 8-12 minutes, and the average duplicate triggers roughly 3 conflicts per year, adding Rp 17,000-Rp 25,000 annually with the midpoint near Rp 21,000.
Layer three distorts acquisition economics. At a blended CAC of Rp 145,000, if 9% of paid-media "new" registrations are actually duplicates of existing customers, effective CAC inflation charges roughly Rp 13,000 per genuine customer back against the duplicate pool — marketing budgets quietly paying twice for people already in the base. Layer four is compliance exposure: under UU No. 27/2022 (the PDP Law), supervisory penalties reach 2% of annual revenue, so amortizing a conservative Rp 500 million expected loss across a 60,000-duplicate base adds about Rp 8,300 per duplicate per year.
Stacked, the retail floor sums storage (about Rp 300), messaging (Rp 2,600 median), reconciliation (Rp 21,000), CAC chargeback (Rp 13,000), and PDP amortization (Rp 8,300) to roughly Rp 45,000. Financial-services stacks carry more weight: eKYC re-verification fees around Rp 3,500 per check through Peruri-connected providers, plus collections teams double-touching households that appear as two separate debtors, lift the cross-sector median to Rp 118,000.
| Cost layer | Annual charge per duplicate | Basis |
|---|---|---|
| Storage | ~Rp 300 | Redundant rows across CRM, POS loyalty, loan/order systems |
| Wasted messaging | Rp 900-Rp 4,300 (median Rp 2,600) | 6-14 touches at Rp 45 SMS / Rp 350 WhatsApp utility |
| Human reconciliation | Rp 17,000-Rp 25,000 (median Rp 21,000) | ~3 conflicts x 8-12 min at Rp 42,000/hr Jakarta agent |
| CAC chargeback | ~Rp 13,000 | 9% duplicate share of paid registrations at Rp 145,000 CAC |
| PDP amortization | ~Rp 8,300 | Rp 500 million expected loss / 60,000 duplicates |
| Retail floor | ~Rp 45,000 | Sum of the five layers above |
| Cross-sector median | ~Rp 118,000 | Adds Rp 3,500/check eKYC re-verification + collections double-touches |
"Duplicates are a cosmetic annoyance — we will clean up someday" is the expensive myth in this ledger. The compound accrues monthly through wasted sends, doubled agent minutes, and inflated CAC, and under the PDP Law's 2%-of-revenue penalty ceiling the do-nothing branch of the 1x10x100 rule has already multiplied the eventual bill a hundredfold by the time a supervisor surfaces it. Before pricing any vendor proposal, rebuild this table from telemetry you already operate: outbound send logs for touch counts, contact-center recordings for conflict minutes, media-platform registration reports for duplicate share, and your legal register for PDP exposure. Four of the five layers are observable without new tooling — and the measured compound, not vendor slideware, is what your 2026 funding case should have to clear.

Benchmark Ledger
USD 12.9 million. According to Gartner's 2021 research note on poor data quality, that is what unreliable records cost the average large organization every year — roughly Rp 204 billion at the Rp 15,800/USD rate prevailing when the note was published. Set that beside the per-record silo tax modeled above and the pattern is structural, not local: duplicated customer identities levy a measurable charge on enterprises everywhere. Indonesia did not invent this tax; it merely priced it.
Two further external bars tighten the frame. According to Experian's Global Data Management Research (2021), 95% of organizations report negative business impact from unreliable data, and the average firm believes 26-29% of its own data is flawed. According to Validity's State of CRM Data Health (2022, built on 1,000+ CRM users surveyed), 44% of respondents doubt their CRM data's accuracy, and respondents estimate roughly 30% of records are stale or duplicated. Together these define the perception bar any internal audit must beat: a stratified sample landing well under that stale-or-duplicated share means you are outrunning the practitioner baseline, while a team insisting its estate is nearly clean is claiming something most CRM operators would not dare assert about their own systems.
The ratio that turns these surveys into sequencing logic comes from Tom Redman's Harvard Business Review article "Seizing Opportunity in Data Quality" (2017): spend Rp 1 verifying a record at capture and you avoid Rp 10 of downstream cleansing and Rp 100 of do-nothing failure cost. That asymmetry is precisely why the decision rule front-loads NIK validation at intake rather than scheduling a grand cleanup. "We will clean up someday" is not a deferral — it is an election. Choosing it books the hundredfold branch of Redman's ratio by default, on an invoice schedule you do not control.
Domestically, the feasibility objection dies on contact with BPJS Kesehatan. According to the agency's 2023 annual report, the JKN database surpassed 270 million registered members running on a single NIK key — one master identity spine spanning nearly the entire population. If a state insurer can hold a register that size coherent, an enterprise estate orders of magnitude smaller has no technical excuse for treating deduplication as impossible. The binding constraint was never infrastructure; it was intake discipline.
Which leaves the enforcement clock. According to Kominfo's supervisory mandate, full PDP Law supervision began in October 2024, with administrative penalties capped at 2% of annual revenue. That date redraws the accounting: 2026 is the first complete budget year in which the silo tax stops being theoretical exposure and becomes an auditable, cash-capable liability. A regulator surfacing your duplicate estate is no longer a hypothetical embarrassment — it is a line item with a statutory ceiling.
| Ledger row | Named source | Verified figure | Decision it forces |
|---|---|---|---|
| Global cost anchor | Gartner research note, 2021 | USD 12.9M average yearly cost (~Rp 204B at Rp 15,800/USD) | Silo tax is worldwide, not an Indonesian anomaly |
| Impact prevalence | Experian Global Data Management Research, 2021 | 95% report negative impact; firms believe 26-29% of data flawed | Expect material flaw rates in any honest audit |
| Perception bar | Validity State of CRM Data Health, 2022 (1,000+ users) | 44% doubt CRM accuracy; ~30% of records stale or duplicated | The baseline your stratified sample must beat |
| Cost-of-delay ratio | Tom Redman, Harvard Business Review, 2017 | Rp 1 at capture : Rp 10 cleansing : Rp 100 do-nothing failure | Front-load NIK validation; cleanup-later elects the x100 branch |
| Feasibility proof | BPJS Kesehatan annual report, 2023 | 270M+ JKN members on a single NIK key | Single-master operation works at national scale in Indonesia |
| Enforcement clock | Kominfo PDP Law supervision, from October 2024 | Penalties capped at 2% of annual revenue | First full auditable budget year is now |
Run the stratified sample against this ledger before the current budget cycle closes. Gartner and Experian bracket the problem's size from above, Redman prices the delay, BPJS Kesehatan removes the feasibility excuse, and Kominfo supplies the deadline — an audit result that beats the Validity perception bar while clearing the program threshold is the strongest funding case the external literature knows how to build.

Four Cleanup Paths, One Winner
Name the four contenders first: (a) manual spreadsheet merges by admin staff; (b) native CRM duplicate rules — Salesforce Duplicate Rules paired with DemandTools; (c) warehouse-native dedup as above; (d) licensed MDM platforms such as Informatica MDM and Ataccama ONE. Path (b) fails quietly for a structural reason worth stating plainly: duplicate rules fire only inside the CRM, so the Tokopedia checkout record and the WhatsApp click-to-chat capture that never entered Salesforce stay invisible to it. That blind spot is why its achievable reduction ceilings at 3–5 points regardless of how well the rules are tuned.
| Path | Year-one cost | Months to payback | Achievable duplicate-rate cut | Ongoing headcount |
|---|---|---|---|---|
| (a) Manual spreadsheet merges | ≈ Rp 120 million (labor) | Never — re-duplication outruns clerical throughput | 2–3 points | Clerical pool that grows with the list |
| (b) CRM-native rules (Salesforce Duplicate Rules + DemandTools) | ≈ Rp 85 million (license + labor) | Usually slower than (c); the 3–5 point ceiling caps recovery | 3–5 points | One CRM admin, part-time |
| (c) Warehouse-native dedup (dbt on BigQuery/Snowflake) | ≈ Rp 520 million (build) | 9–14 months | 5–8 points | Small analytics-engineering pod |
| (d) Licensed MDM (Informatica MDM, Ataccama ONE) | ≈ Rp 2.4 billion (license + integration) | Viable only above roughly 5 million records | 6–9 points | Dedicated stewardship team plus vendor support |
Read row (a) twice, because it is the trap. At roughly Rp 120 million with no license fee, manual merging presents itself as the fiscally responsible default — which is precisely how the "duplicates are cosmetic, we will clean up someday" reflex survives procurement review after procurement review. What kills it is arithmetic, not opinion: clerical throughput scales linearly with headcount, while the duplicate population scales with list growth. At 8% annual record growth, re-duplication outruns the merging team, the backlog regrows faster than it is cleared, and the path never reaches payback — the only option on the board with an infinite payback period. Meanwhile the per-record tax quantified at the top of this guide accrues monthly, so the do-later branch keeps compounding while the spreadsheet team treads water.
One criterion separates winners from losers and appears in none of the cost columns: merge lineage. Only paths (c) and (d) produce an auditable survivorship log — which record survived, which was retired, which rule decided, who approved it. If Kominfo requests data provenance during a PDP Law supervision review, that log is the artifact you hand over; paths (a) and (b) leave no traceable record of who merged what, converting a routine supervision question into an unplanned forensic project. With dbt the lineage arrives almost free, since every merge is a version-controlled SQL model with a full change history, whereas MDM resells the same traceability behind certification overhead. That traceability has become valuable enough that niche advisories now build whole frameworks around it — Glacial Wolf Strategy's 2026 material packages exactly this conversion of scattered governance artifacts into traceable decisions — a useful signal of where audit expectations are heading.
Two boundary conditions finish the map. Above roughly 5 million records, or whenever three or more domains — customer, product, supplier — need governed golden records, MDM's Rp 2.4 billion stops being overkill and becomes the only architecture that holds. Below both thresholds in the decision rule above, buy nothing at all: enforce NIK/NPWP keys at every intake point and let quarterly rule-based merges inside the existing CRM carry the load. And before sitting through any vendor demo, run the lineage test: ask to see a survivorship log entry from the demo tenant showing who merged what, when, and under which rule. If the sales engineer cannot produce one, neither can the platform.

What the Data Doesn't Tell You
Merriam-Webster's usage guidance, revised August 17, 2026, defines "disconnected" as "not connected: separate; also: incoherent." Both senses describe the evidence base behind the silo-tax model. The headline figure is a calibrated midpoint assembled from allocation assumptions — agent reconciliation minutes, wasted messaging volume, inflated acquisition cost — none of which appears on any invoice labeled "duplicate." Two enterprises with identical duplicate rates can realize materially different losses depending on how their finance teams attribute shared labor, so the model prices a mechanism, not a meter reading.
The measurement layer has its own fragility. Match quality depends heavily on Indonesian naming conventions: a large share of citizens hold a single legal name, leaving probabilistic matchers no surname signal to exploit, and transliterated Chinese-Indonesian names compound the ambiguity. Measured duplicate rates therefore carry error bars that are widest exactly where the decision rule's five-percent cutoff sits. A borderline stratified sample is not a verdict; it is a prompt to draw a second, independently stratified sample before committing budget.
Variance across cases is structural, not noise. In consumer lending, each duplicate tends to trigger its own KYC re-verification, another inquiry into Pefindo's SLIK credit-information system, and collections routed to two addresses for one debtor — which is why the modeled premium concentrates at the top of the spread described earlier. Retail-floor damage skews toward redundant SMS and WhatsApp sends and split loyalty baskets, pricing far lower. Channel mix moves the number as well: phone-only click-to-chat capture leaves fewer match keys than email checkout, so two retailers with equal duplicate rates can face unequal remediation effort.
Three conditions strain the rule without overturning it. First, mandated remediation: when an OJK examination or a PDP Law audit surfaces duplication before management does, the do-nothing branch has been repriced by the supervisor, and payback arithmetic stops being the governing frame — the rule prices discretionary spend, not compulsory correction. Second, non-stationary record bases: an acquisition that will lift a company past the record-count threshold within a few quarters argues for re-running the stratified sample at integration rather than treating today's count as permanent. Third, cadence mismatch: the below-threshold prescription of NIK/NPWP keys with quarterly rule-based merges assumes batch tolerance; a real-time personalization stack may need a shorter merge cycle inside the same CRM — an operational tightening, not a platform purchase.
| Edge case | What distorts | Correct response | Rule status |
|---|---|---|---|
| Sample lands near the 5% cutoff | Sampling error can straddle the line | Draw a second independent stratified sample | Holds — measure-first logic |
| Mononym-heavy customer base | Fuzzy name matching loses its main signal | Weight NIK and phone keys over name similarity | Holds |
| Pending acquisition or migration | Record count treated as static | Re-sample at integration, then decide | Holds — timing shifts |
| OJK-supervised lender flagged by examiner | Remediation becomes mandatory | Treat as a compliance project; ROI secondary | ROI branch inapplicable |
| Real-time personalization stack | Quarterly merge window too coarse | Shorten the merge cycle within the CRM | Holds — cadence, not platform |
The stubborn counter-argument — duplicates are a cosmetic annoyance, we will clean up someday — fails on timing, not magnitude. The tax accrues monthly through wasted sends, doubled reconciliation minutes, and inflated CAC, and by the time a regulator or auditor surfaces it under the PDP Law, the do-nothing branch of the 1x10x100 rule has already run its course. Deferral is not a neutral default; it is the priciest option on the menu, simply the one whose invoice arrives last.
If a stratified sample lands near the cutoff, spend the next hour auditing the sample design itself: confirm the strata reflect channel mix and region, check whether mononyms were scored as low-confidence matches or discarded outright, and document the false-match rate before anyone rounds the percentage upward in a budget meeting. The model is robust enough to fund a program; it is not precise enough to skip the measurement discipline that produced it.

Where the Model Breaks
Stress the silo-tax model against Indonesian data and it bends in six places — one hard enough to flip a purchase verdict from buy to don't-buy. None of the bends rescue the do-nothing option, but all six change what a defensible business case looks like. Start with dormancy: audits at Indonesian retailers typically find that only 30–40% of duplicate pairs were messaged or transacted within the trailing twelve months, while the rest sit inert. Pricing every twin at the full Rp 118,000 compound modeled earlier therefore overstates realized loss by roughly 2.5 times, because a dormant duplicate generates no wasted WhatsApp sends, no doubled agent reconciliation minutes, and no inflated acquisition cost until it wakes up. Tier the tax by last-activity recency; even the deflated figure leaves the funding thresholds intact for qualifying books.
Second, the matcher moves the denominator. Switching from exact phone and NIK keying to Levenshtein fuzzy name matching swings detected duplicate counts by plus-or-minus 40% on Indonesian name pools, because patronymic staples such as Budi Santoso or Agus Wijaya spawn false-positive clusters that inflate the tax estimate and the promised savings in the same direction. Any vendor quoting a single duplicate count without naming its matcher and threshold is selling a confidence interval collapsed into a point. Run both passes, publish both bounds, and fund against the exact-key floor.
Third, interrogate the baseline inside vendor math. According to Forrester's Total Economic Impact studies commissioned by Microsoft — which produced the only concrete payback durations in the available record, six months for Business Central and 16 months for a midmarket Dynamics 365 ERP deployment — the modeled environments assume starting duplicate rates of 15–25%. Indonesian banks that have enforced NIK capture at onboarding since 2020 frequently begin at 2–4%, a starting position that stretches modeled payback past 36 months and, per the decision rule above, legitimately flips the verdict to don't-buy. Your sampled rate, not the template's, sets the clock.
Fourth, the regulatory line item. The PDP Law's fine ceiling — up to two percent of annual revenue — behaves like an accrual on paper, but enforcement requires regulators to prove causation between duplication and a reported violation. Booking that ceiling as an annual operating cost rather than a low-probability tail injects roughly Rp 30,000 or more of phantom cost per record into the model. Book it as probability-weighted exposure, disclosed separately, or the model stops being an estimate and becomes advocacy.
Fifth, event risk. Core-system migrations and post-acquisition customer-book consolidations reset duplicate baselines overnight: merging two legacy CRMs into one warehouse reintroduces at scale the exact fragmentation the audit just cleared. A spotless review closed before an integration lands in 2026 predicts nothing about exposure after go-live, so the stratified sample belongs in the cutover checklist itself, re-run immediately post-migration rather than annually.
Sixth, survivorship. Dedup programs that die usually die at the branch-manager level — front-line leaders who refuse merged customer views because single-customer visibility exposes portfolio gaming and inter-branch poaching — and those failures almost never reach vendor case libraries, which is why observed success rates are structurally overstated. Reporting from Relic.inc in April 2026 makes the adjacent point: technology-first deployments without a defined business problem end as expensive tools with no measurable link to value.
| Stress test | What the data shows | Effect on the model | Corrective |
|---|---|---|---|
| Dormancy screen | Only 30–40% of duplicates active in trailing 12 months | Realized loss overstated ~2.5x | Tier tax by last-activity recency |
| Matcher swap | Exact phone/NIK vs Levenshtein fuzzy names | Detected counts swing ±40% | Publish both bounds; fund on exact-key floor |
| Baseline realism | TEI templates assume 15–25%; NIK-enforced banks start 2–4% | Payback stretches past 36 months | Re-model on your own sampled rate |
| Penalty treatment | PDP fines require proven causation | Roughly Rp 30,000+/record phantom cost | Book as probability-weighted tail |
| Cutover events | Migrations and acquisitions reset baselines overnight | Pre-integration audits lose predictive power | Re-sample immediately post-go-live |
| Survivorship | Branch-resistance failures absent from vendor case libraries | Success rates structurally overstated | Demand references from killed projects |
The comfortable myth that duplicates are a cosmetic annoyance to clean up someday survives partly because the counter-evidence is unpublished — but the tax accrues monthly on the active slice of the book whether or not anyone models it precisely. Sample twice, once on exact NIK and phone keys and once on fuzzy names, attach a recency-tiered loss estimate to both bounds, and let the decision rule's thresholds — not a vendor template — decide whether a platform gets funded.

Worked Case
Month ten or month seventeen — that spread is the entire purchase decision, and one worked case shows where the hinge sits. Take a composite but fully specified subject: a Bandung-headquartered omnichannel retailer holding 420,000 active customer records split between its POS loyalty database and its webshop CRM. A stratified audit of that base — sampled by channel, tenure, and recency, per the decision rule, not pulled as one convenience file — returns a 7.4% duplicate rate: 31,080 records that exist twice. Every input below is stipulated so you can rerun the arithmetic against your own audit extract; nothing depends on trusting the case.
Apply the retail cost stack of Rp 74,000 per duplicate-year to those 31,080 records and the modeled drag is Rp 2.30 billion annually. Nobody banks that number, and pretending otherwise is how deduplication programs die in their first quarterly review. The honest step is the realization haircut: only active duplicates get addressed, and channel coverage is partial, so a 30% realization factor cuts realistic annual recovery to Rp 690 million. Business cases that book the full modeled figure set themselves up to be falsified by their own delivery data.
Cost the fix with the same discipline. A warehouse-native build — dbt matching models on BigQuery, no new platform license — runs Rp 520 million one-time plus Rp 60 million per year to operate. Year-one net position: Rp 690 million minus Rp 60 million minus Rp 520 million, or Rp 110 million positive. Payback is the Rp 520 million build divided by Rp 630 million in net annual recovery: roughly month ten, squarely inside the 9–14 month window.
Now falsify it. Hold everything constant — same 420,000 records, same stack, same realization factor, same build — and drop the duplicate rate to 4%. Because record volume is fixed, the rate is the only moving part, which makes this a clean isolation of the threshold. The result: 16,800 duplicates, Rp 1.24 billion in modeled drag, Rp 373 million recovered. Breakeven slides to roughly month seventeen on a gross-recovery basis, nearer month twenty once the run cost is charged against payback — outside the window on either accounting convention.
That gap is why the decision rule's 5% floor is load-bearing rather than decorative. At 7.4%, this retailer funds a dedicated program and recovers the build inside year one. At 4.0% — still 16,800 polluted records carrying Rp 1.24 billion of modeled drag — the rule routes the same company to the fallback: enforce NIK/NPWP keys at every intake point and run quarterly rule-based merges inside the existing CRM, buying no platform. Note what the fallback is not. It is not "clean up someday." The tax accrues monthly in both branches through wasted sends and doubled reconciliation minutes, which is exactly why the sub-threshold branch gets intake controls immediately instead of a deferred cleanup promise.
| Ledger line | Audit case (7.4%) | Falsified case (4.0%) |
|---|---|---|
| Active records | 420,000 | 420,000 |
| Duplicate records | 31,080 | 16,800 |
| Modeled annual drag (Rp 74,000/duplicate-year) | Rp 2.30 billion | Rp 1.24 billion |
| Realized recovery (30% factor) | Rp 690 million/yr | Rp 373 million/yr |
| Build + run cost | Rp 520 M + Rp 60 M/yr | Rp 520 M + Rp 60 M/yr |
| Year-one net position | +Rp 110 million | −Rp 207 million |
| Breakeven and verdict | Month 10 — fund the program | Month 17+ — take the fallback |
Replicate this ledger on your own base before any vendor conversation this year: pull this quarter's stratified duplicate rate, multiply by your sector's per-duplicate stack from the benchmark ledger above, apply a 30% realization factor, then divide build cost by net annual recovery. Breakeven inside fourteen months funds a program; breakeven past it means the money belongs in intake keys, not a platform.
How to Choose Well
Most deduplication programs fail on the funding decision, not the matching logic. The five rules below form a single decision tree: pass a gate, descend to the next; fail one, take the fallback branch and spend nothing. Applied in order through 2026 — with PDP Law enforcement now routine enough that auditors ask about customer-record hygiene unprompted — they keep the purchase decision honest.
The tree exists to kill one belief outright: that duplicates are a cosmetic annoyance to clean up someday. The tax accrues monthly through wasted SMS and WhatsApp sends, doubled agent reconciliation minutes, and inflated acquisition costs — and once a regulator or auditor surfaces it under the PDP Law, the do-nothing branch of the 1×10×100 rule has already multiplied the eventual bill a hundredfold. "Someday" is the expensive branch.
Rule 2 — Prevent at the key. Before spending anything on cleanup tooling, mandate NIK validation through the Dukcapil check for individuals and NPWP-16 for businesses at every intake form. The 1×10×100 ratio makes entry-point verification roughly ten times cheaper than downstream repair. Regulated industries already run this way: according to MasterControl, each system in a regulated environment must create a production record for every batch, lot, or unit before market release — identity keys deserve the same upstream discipline.
Rule 3 — Match tool to scale. Warehouse-native deduplication — dbt running on BigQuery or Snowflake — covers 150,000 to 1,000,000 active records. Reserve Informatica- or Ataccama-class MDM platforms for 5 million-plus records or three-plus domains that need governed golden records. The unmapped band between 1 million and 5 million single-domain records is deliberate judgment territory: stay warehouse-native until cross-domain governance, not raw record count, forces the upgrade.
Rule 4 — Underwrite on realized capture. Build every business case at 25–40% of the modeled drag, never above 50%. Two mechanisms justify the haircut: duplicates regenerate at intake until Rule 2 lands, and fuzzy matching always carries false positives whose merges destroy good records alongside bad ones. Reject outright any proposal promising recovery of the full modeled silo tax — that figure is a ceiling on the problem, not a forecast of savings.
| # | Gate | Action if passed | Fallback if failed |
|---|---|---|---|
| 1 | Audited rate on a stratified 10,000-record sample, two matchers | Fund a dedicated program above 5% | Quarterly rule-based merges in the existing CRM; revisit next cycle |
| 2 | Every intake form, before cleanup tooling | Dukcapil NIK check (individuals); NPWP-16 (businesses) | Unverified records enter pre-duplicated |
| 3 | Active-record count and domain count | dbt on BigQuery or Snowflake for 150,000–1,000,000 records | Informatica- or Ataccama-class MDM at 5M+ records or 3+ domains |
| 4 | Business-case review | Underwrite at 25–40% of modeled drag, 50% hard ceiling | Reject any full-recovery promise outright |
| 5 | Contract signature | Kill-switch: audited rate below 3% within 12 months of go-live | Halt run-rate spend; revert to quarterly manual merges |
Apply the gates in sequence and let the first failed gate set the budget — everything after it is negotiation, not analysis.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Pull a stratified sample of active records across the CRM, POS loyalty database, and loan/order system — strata by intake channel (Tokopedia checkout, WhatsApp click-to-chat, POS enrollment) — and match on normalized phone, email, and name to compute your true duplicate rate. | The whole funding decision hangs on this measurement; the silo tax (Rp 45,000 retail floor, Rp 118,000 financial-services median) cannot be sized or approved until the duplicate share of active records is measured, not assumed. |
| 2 | Apply the gate: fund a dedicated deduplication program only if duplicates exceed 5% of active records AND you hold 150,000+ records. Below either threshold, skip procurement entirely. | Under those conditions the correct stack is NIK/NPWP enforcement plus quarterly rule-based merges inside the existing CRM — buying an MDM platform there costs more than the tax it removes. |
| 3 | If you clear both thresholds, size the dedicated program against the stacked per-duplicate charge and approve only on a payback point inside 16 months with a ~200% return. | At Rp 45,000 per duplicate in retail or Rp 118,000 in financial services, a program that hits payback within 16 months self-funds; anything slower signals scope creep, not savings. |
| 4 | Close the key vacuum at every intake point: require the NIK (NPWP for business entities) before a row is minted in the CRM, POS loyalty database, and loan/order system — including the Tokopedia checkout handoff and the WhatsApp click-to-chat form that currently captures only a phone number. | One human becoming three billing entities is an architecture failure, not staff carelessness; enforcing the state-issued identifier stops new duplicates at the source so the tax fades instead of recurring. |
| 5 | On the below-threshold path, run quarterly rule-based merges inside the existing CRM: exact NIK matches first, then normalized phone + name and email + name rules, collapsing each cluster into a single survivor record. | Deterministic rules alone typically resolve around 38% of duplicate pairs, recovering the messaging waste (Rp 900–Rp 4,300 per duplicate) and the ~Rp 21,000 annual reconciliation load without a new platform. |
| 6 | Log every merge decision and NIK access event in an auditable trail aligned to UU No. 27/2022 (PDP Law), and route high-stakes survivors — especially in the loan/order system — through Peruri-connected eKYC re-verification. | Supervisory penalties reach 2% of annual revenue, the exposure already running ~Rp 8,300 per duplicate per year; ungoverned merging trades a messaging bill for a compliance one. |
Quick answers
| What is the modeled annual silo tax per duplicate customer record in Indonesia? | Roughly Rp 45,000 per duplicate per year for retail stacks and Rp 118,000 as the cross-sector median once eKYC re-verification fees around Rp 3,500 per check and collections double-touches are added. |
| How much does wasted messaging contribute per duplicate each year? | Rp 900-Rp 4,300 annually (median Rp 2,600), driven by 6-14 redundant outbound touches at Telkomsel's enterprise SMS rate of roughly Rp 45 per message and WhatsApp Business API utility-template rates near Rp 350. |
| What does human reconciliation of conflicting records cost? | Rp 17,000-Rp 25,000 per duplicate annually with the midpoint near Rp 21,000, based on roughly 3 conflicts per year taking 8-12 minutes each for a fully loaded Jakarta contact-center agent costing about Rp 42,000 per hour. |
| What is the payback logic that justifies front-loading NIK validation at intake? | Tom Redman's 2017 Harvard Business Review 1x10x100 rule: spend Rp 1 verifying a record at capture and you avoid Rp 10 of downstream cleansing and Rp 100 of do-nothing failure cost. |
| Why does 2026 mark the payback turning point for fixing duplicates? | Because Kominfo's supervisory mandate started full PDP Law supervision in October 2024 with administrative penalties capped at 2% of annual revenue, 2026 is the first complete budget year in which the silo tax stops being theoretical exposure and becomes an auditable, cash-capable liability. |