---
name: document-set-extraction-table
description: Extracts a defined field set from every same-kind document in a folder into a structured sheet with source citations, confidence flags, and a separate exceptions register. Use for batch contract, invoice, CV, or policy extraction, document-set coding, field harvesting, and cited folder-wide data capture.
license: Apache-2.0
metadata:
  adlass.categories: "data-tables/enrichment, documents-contracts/document-extraction"
  adlass.industries: ""
  adlass.tags: "batch-extraction, document-set, field-capture, citations, confidence, exceptions"
  adlass.adaptation: "mapping"
  adlass.source: "original"
  adlass.version: "1"
---

# Document-set extraction table

## Purpose

Convert a folder of same-kind documents into one auditable extraction sheet. Each document becomes one row, each requested field becomes a value with a precise citation and a confidence flag, and unresolved cases are listed separately for review.

## Scope

Use for a defined collection such as supplier contracts, invoices, CVs, insurance certificates, or policies where the same field schema applies to every document. The process identifies document coverage, maps values to the requested columns, preserves units and dates, and records contradictions or missing evidence.

**Excluded:** interpreting legal or commercial meaning beyond the requested fields, inventing values from context, normalising data into an unprovided company master schema, and changing the source documents.

## Data basis

- Every document in the folder supplied to the run, treated as one member of the declared document set.
- The `document_kind` input describing the expected kind of document.
- The `requested_fields` input defining field names, value formats, and any allowed values.

## Result

An extraction sheet with one row per in-scope document and field columns containing the value, citation, and confidence flag. An exceptions sheet lists omitted documents, missing fields, conflicting values, unreadable passages, schema violations, and assumptions.

## Quality criteria

- The coverage register accounts for every folder document exactly once, including documents excluded as the wrong kind.
- Every populated field has a document identifier and a section, page, table, or row citation.
- Missing, conflicting, and unreadable fields use explicit status values rather than blank cells.
- Confidence is exactly `high`, `medium`, or `low`, with a reason for every `medium` or `low` value.
- Dates, amounts, currencies, names, and identifiers retain source precision; conversions are recorded in the exceptions sheet.
- The extraction sheet has one stable row key per source document and no duplicate document rows.

## Instructions

Treat the requested field list as the schema: do not add inferred columns or silently rename fields. Use the document's own labels and context to resolve a value, but never select between contradictory values without recording both citations. Mark a field `not stated` when the document does not provide it, `unclear` when the passage cannot be resolved, and `not applicable` only when the document explicitly makes the field inapplicable. A citation must be specific enough for a reviewer to find the evidence without searching the whole folder. Keep source formatting in the cited value and put any normalized representation in a separate normalized field only when the input schema asks for it. Reconcile the row count against the folder inventory before finishing.

## Adapt before use

- Define the document kind and the exact field schema, including required fields, formats, and allowed values, in the run input.
- Map source labels and normalized output columns for the company's terminology, identifiers, date formats, and currency handling.
- Decide whether a document with a wrong kind, duplicate file, or unreadable content belongs in the extraction sheet, the exceptions sheet, or both.
- Set the confidence policy and the roles who review low-confidence or contradictory fields.
