Volver a la biblioteca

Record deduplication plan

Analyzes a record table for duplicate entities, combines exact and fuzzy identity evidence into proposed merge groups, selects a surviving record, and defines field-level merge rules in a reviewable plan. Use for customer, supplier, contact, account, master-data, duplicate-record, and CRM deduplication work.

por adlass TemplatesVersión 1Usa herramientas de adlassUniversal

Publicado 21 de ago de 2026 · Actualizado 26 de ago de 2026

Útil · 0Ver SKILL.md sin formato

Requisitos

Declare the entity type, stable ID column, identity columns, and unique identifiers in match_policy. Set normalization rules and fuzzy thresholds for every comparison field used by the table. Define survivor precedence from verification, recency, lifecycle, and completeness fields. Specify fields that may be combined and the reviewer role for manual decisions.

Documento de la habilidad

El SKILL.md completo que tu agente lee y sigue.

Record deduplication plan

Purpose

Turn a record table into a conservative, reviewable deduplication plan. The skill identifies records that may represent the same customer, supplier, contact, or account, explains the exact and fuzzy evidence, chooses a proposed surviving record for each group, and specifies how each field should be retained, combined, or escalated.

Scope

Use on one structured record table with a stable record identifier and entity attributes such as name, email, phone, address, tax or registration ID, website, and account status. The plan covers normalization for matching, exact-key collisions, fuzzy candidate generation, match grouping, survivor selection, field-level merge decisions, and unresolved ambiguity.

Excluded: deleting or updating source records, executing merges, contacting entities, enriching records from the web, and deciding ownership or legal identity where the table lacks evidence.

Data basis

  • The complete record table in the run scope, including its schema, row identifiers, values, nulls, and data types.
  • The entity type and identity rules supplied in the entity_type and match_policy inputs.
  • Record history, timestamps, status, and relationship counts when those columns exist in the table.

Result

A merge-plan sheet with one row per proposed duplicate pair or group, exact and fuzzy evidence, confidence, surviving record, and field-level actions. A review memo reports coverage, match decisions, high-risk groups, and records intentionally left separate.

Quality criteria

  • Every input row is counted once in the coverage summary and has either no candidate, a proposed group, or an unresolved exception.
  • Exact matches identify the normalized key and source column; fuzzy matches identify the fields, similarity rationale, and conflicting evidence.
  • No record belongs to two proposed groups, and every group has a stable group ID and a cited row identifier for each member.
  • Each group names one surviving record with a concrete selection reason or is marked unresolved.
  • Field-level rules use only keep_survivor, fill_blank, choose_newest, combine_distinct, or manual_review.
  • No merge is proposed for weak evidence, conflicting unique identifiers, or incompatible entity types.

Instructions

Normalize only for comparison: trim whitespace, case-fold text, standardize punctuation, and apply the declared handling for phone, email, website, address, and identifier formats. Preserve original values in the plan. Treat an exact match on a declared unique identifier as strong evidence, but do not merge rows when another identity field materially conflicts. Use fuzzy similarity to create candidates, not as proof by itself; explain which fields agree and which disagree. Prefer a surviving row with a stable ID, complete identity fields, current status, and the most recent verified activity, in that order unless match_policy states otherwise. Never overwrite a nonblank survivor value with a blank or unverified value. When two nonblank values cannot be reconciled, use manual_review and list both source values. Record the evidence row IDs and original column names for every group decision.

Adapt before use

  • Declare the entity type, stable ID column, identity columns, and which identifiers are unique in match_policy.
  • Set normalization rules and fuzzy-match thresholds for names, emails, phones, addresses, websites, and local identifiers.
  • Define survivor precedence using the company's verification, recency, lifecycle, and completeness fields.
  • Specify fields that may be combined, fields that must remain single-valued, and the reviewer role for manual decisions.

Habilidades relacionadas

  • Sync a spreadsheet into a record type

    Loads an uploaded spreadsheet into a record type of the workspace with the bulk sync tool, previewing the changes before writing. Use for master data imports, periodic refreshes of a table, or migrating a list from Excel into records.

  • Anomaly variance explanation

    Explains material changes in a business metric by decomposing period, volume, price, mix, and data-quality effects, then ranks evidence-backed causes in a cited analysis sheet and management memo. Use for KPI variance reviews, monthly business reviews, forecast misses, anomaly investigation, and driver analysis.

  • Data migration field mapping

    Maps source-system fields to a target data model for migration, documenting transformations, requiredness, value translations, identifiers, validation rules, and unresolved collisions. Use for CRM migrations, ERP imports, warehouse loads, application replacement, and controlled spreadsheet-to-system field mapping.

  • Dataset Quality Assessment

    Dataset Quality Assessment converts table schema, column types, null policy, primary and foreign keys into profiling report and defect register with field metrics, duplicate groups, invalid values, orphan keys, temporal gaps, and severity, with cited evidence, explicit rules, and unresolved exceptions. Use for dataset quality assessment, recurring review, and decision preparation.