---
name: document-classification-register
description: Classifies each document against a supplied taxonomy and records document type, business purpose, sensitivity, retention cue, owner, and rationale. Use for records inventories, contract libraries, information-governance reviews, and knowledge-base cleanup.
license: Apache-2.0
metadata:
  adlass.categories: "documents-contracts/classification"
  adlass.industries: ""
  adlass.tags: "classification,taxonomy,records,sensitivity,retention,document-inventory"
  adlass.adaptation: "reference-doc"
  adlass.source: "original"
  adlass.version: "1"
---

# Document classification register

## Purpose

Turn a document corpus into a consistent classification register that explains what each file is, why it matters, how sensitive it is, and which retention cue applies.

## Scope

Classify filename, MIME type, title, author, dates, owner, business process, document type, confidentiality, personal-data indicator, retention category, version, and duplicate relationship. Use the supplied taxonomy and records policy.

**Excluded:** deleting or renaming files, changing permissions, legal retention advice, and inventing an owner or retention period.

## Data basis

- Document corpus with file ID, path, filename, MIME type, title, created and modified dates, and readable text.
- Classification taxonomy, sensitivity labels, records schedule, retention policy, and owner directory.
- Optional existing inventory and duplicate hashes.

## Result

A persistent register with one row per document, assigned taxonomy path, purpose, sensitivity, personal-data flag, retention cue, owner, confidence, rationale, and source citation.

## Quality criteria

- Every source file has exactly one register row or an explicit unreadable/excluded status.
- Assigned labels are valid taxonomy values and rationale cites filename, metadata, or text evidence.
- Similar files are linked without collapsing distinct versions.
- Low-confidence and policy gaps are visible.

## Instructions

Use metadata as evidence only for metadata fields; use document text for purpose and sensitivity cues. Prefer the most specific taxonomy node supported by evidence. Do not treat a filename keyword as proof of personal data. Keep duplicates as separate rows and link their IDs. Apply retention cues, not invented retention durations, when the policy lacks a matching category.

Record the document language and version marker when present because both can affect classification confidence. A policy match should identify the schedule category, not merely repeat the file’s title. Keep classification rationale short enough to audit row by row.

## Adapt before use

- Add the document taxonomy, sensitivity labels, records schedule, and owner directory.
- Define duplicate handling, unreadable-file status, and the confidence vocabulary.
- Map file metadata fields and approved business-process names.
