---
name: record-deduplication-plan
description: Analyzes a record table for duplicate entities, combines exact and fuzzy identity evidence into proposed merge groups, selects a surviving record, and defines field-level merge rules in a reviewable plan. Use for customer, supplier, contact, account, master-data, duplicate-record, and CRM deduplication work.
license: Apache-2.0
metadata:
  adlass.categories: "data-tables/cleanup-normalization"
  adlass.industries: ""
  adlass.tags: "deduplication, duplicate-records, entity-resolution, merge-plan, master-data, fuzzy-match"
  adlass.adaptation: "mapping"
  adlass.source: "original"
  adlass.version: "1"
---

# Record deduplication plan

## Purpose

Turn a record table into a conservative, reviewable deduplication plan. The skill identifies records that may represent the same customer, supplier, contact, or account, explains the exact and fuzzy evidence, chooses a proposed surviving record for each group, and specifies how each field should be retained, combined, or escalated.

## Scope

Use on one structured record table with a stable record identifier and entity attributes such as name, email, phone, address, tax or registration ID, website, and account status. The plan covers normalization for matching, exact-key collisions, fuzzy candidate generation, match grouping, survivor selection, field-level merge decisions, and unresolved ambiguity.

**Excluded:** deleting or updating source records, executing merges, contacting entities, enriching records from the web, and deciding ownership or legal identity where the table lacks evidence.

## Data basis

- The complete record table in the run scope, including its schema, row identifiers, values, nulls, and data types.
- The entity type and identity rules supplied in the `entity_type` and `match_policy` inputs.
- Record history, timestamps, status, and relationship counts when those columns exist in the table.

## Result

A merge-plan sheet with one row per proposed duplicate pair or group, exact and fuzzy evidence, confidence, surviving record, and field-level actions. A review memo reports coverage, match decisions, high-risk groups, and records intentionally left separate.

## Quality criteria

- Every input row is counted once in the coverage summary and has either no candidate, a proposed group, or an unresolved exception.
- Exact matches identify the normalized key and source column; fuzzy matches identify the fields, similarity rationale, and conflicting evidence.
- No record belongs to two proposed groups, and every group has a stable group ID and a cited row identifier for each member.
- Each group names one surviving record with a concrete selection reason or is marked unresolved.
- Field-level rules use only `keep_survivor`, `fill_blank`, `choose_newest`, `combine_distinct`, or `manual_review`.
- No merge is proposed for weak evidence, conflicting unique identifiers, or incompatible entity types.

## Instructions

Normalize only for comparison: trim whitespace, case-fold text, standardize punctuation, and apply the declared handling for phone, email, website, address, and identifier formats. Preserve original values in the plan. Treat an exact match on a declared unique identifier as strong evidence, but do not merge rows when another identity field materially conflicts. Use fuzzy similarity to create candidates, not as proof by itself; explain which fields agree and which disagree. Prefer a surviving row with a stable ID, complete identity fields, current status, and the most recent verified activity, in that order unless match_policy states otherwise. Never overwrite a nonblank survivor value with a blank or unverified value. When two nonblank values cannot be reconciled, use `manual_review` and list both source values. Record the evidence row IDs and original column names for every group decision.

## Adapt before use

- Declare the entity type, stable ID column, identity columns, and which identifiers are unique in `match_policy`.
- Set normalization rules and fuzzy-match thresholds for names, emails, phones, addresses, websites, and local identifiers.
- Define survivor precedence using the company's verification, recency, lifecycle, and completeness fields.
- Specify fields that may be combined, fields that must remain single-valued, and the reviewer role for manual decisions.
