Dataset Quality Assessment
Dataset Quality Assessment converts table schema, column types, null policy, primary and foreign keys into profiling report and defect register with field metrics, duplicate groups, invalid values, orphan keys, temporal gaps, and severity, with cited evidence, explicit rules, and unresolved exceptions. Use for dataset quality assessment, recurring review, and decision preparation.
Published Aug 21, 2026 · Updated Aug 26, 2026
Requirements
Map the process identifiers, source columns, and row citation convention. Add the governing policy, rubric, schema, or reference tables to the scope. Define thresholds, status values, units, and exception severity.
Skill document
The full SKILL.md your agent reads and follows.
Dataset Quality Assessment
Purpose
Dataset Quality Assessment turns the supplied records into profiling report and defect register with field metrics, duplicate groups, invalid values, orphan keys, temporal gaps, and severity. It keeps source facts, derived values, recommendations, and limitations separate so each conclusion can be checked.
Scope
Cover table schema, column types, null policy, primary and foreign keys, allowed values, row counts, dates, freshness, lookup tables. Preserve source identifiers, dates, units, and wording needed to trace every result.
Excluded for dataset fitness and defect triage: changing source records, taking external actions, making a human approval or employment decision, and asserting facts absent from the corpus.
Data basis
- Process documents and tables containing table schema, column types, null policy, primary and foreign keys, allowed values, row counts, dates, freshness, lookup tables.
- Company policy, rubric, schema, templates, and reference tables in scope.
- Run input for period, audience, entity, or threshold when it changes this run.
Result
Produce profiling report and defect register with field metrics, duplicate groups, invalid values, orphan keys, temporal gaps, and severity. Each material row and conclusion cites a document, section, page, row, field, or run input.
Quality criteria
- Every row-level quality metrics and rule coverage record is included or has an exclusion reason.
- Calculations state fields, units, denominator, and rule.
- Missing, contradictory, stale, and ambiguous values stay labelled rather than guessed.
- Facts, interpretations, and proposed next actions use separate fields.
- Supplied terminology, thresholds, and status values are used consistently.
- The final limitations section states what the corpus could not establish.
Instructions
Use the data dictionary and quality policy. Profile all indexed rows with denominators. Distinguish null, blank, malformed, duplicate, out-of-domain, and orphan values; cite row keys and rule IDs. Preserve original values beside normalized values, cite every material finding, and use “not assessed” when required evidence or a rule is absent. When sources disagree, show both citations and explain the conflict. Keep row-level evidence in the supporting sheet and summarize only supported conclusions in the document.
Adapt before use
- Map the schema and primary-key mapping fields and row citation convention.
- Add the governing policy, rubric, schema, or reference tables to the scope.
- Define thresholds, status values, units, and exception severity.
Related skills
- Sync a spreadsheet into a record type
Loads an uploaded spreadsheet into a record type of the workspace with the bulk sync tool, previewing the changes before writing. Use for master data imports, periodic refreshes of a table, or migrating a list from Excel into records.
- Anomaly variance explanation
Explains material changes in a business metric by decomposing period, volume, price, mix, and data-quality effects, then ranks evidence-backed causes in a cited analysis sheet and management memo. Use for KPI variance reviews, monthly business reviews, forecast misses, anomaly investigation, and driver analysis.
- Data migration field mapping
Maps source-system fields to a target data model for migration, documenting transformations, requiredness, value translations, identifiers, validation rules, and unresolved collisions. Use for CRM migrations, ERP imports, warehouse loads, application replacement, and controlled spreadsheet-to-system field mapping.
- Document-set extraction table
Extracts a defined field set from every same-kind document in a folder into a structured sheet with source citations, confidence flags, and a separate exceptions register. Use for batch contract, invoice, CV, or policy extraction, document-set coding, field harvesting, and cited folder-wide data capture.