Back to library

Document field extraction

Extracts a defined schema from mixed business documents into a normalized, source-cited table while preserving original wording, missing fields, and confidence notes. Use for invoices, forms, receipts, applications, statements, certificates, and document registers.

by adlass TemplatesVersion 1Uses adlass toolsUniversal

Published Aug 21, 2026 · Updated Aug 26, 2026

Helpful · 0View raw SKILL.md

Requirements

Add the field dictionary with types, required flags, aliases, and allowed values. Define source identifiers, repeating-field layout, date, number, currency, and missing-value rules.

Skill document

The full SKILL.md your agent reads and follows.

Document field extraction

Purpose

Turn a heterogeneous document set into one row per source document and one column per declared field. Preserve the raw value, normalized value, source location, and extraction note so downstream users can audit a date, amount, identifier, party name, or line-item total without reopening the entire corpus.

Scope

Extract only the schema declared in the mapping document or run input. Handle repeated labels, tables, headers, footers, multi-page records, handwritten or unreadable values, and fields that are absent. Keep document-level fields separate from repeating line items and never merge two source documents because their names appear similar.

Excluded: classifying documents into business categories, deciding whether a value is legally valid, correcting a source record, and filling a blank from external knowledge.

Data basis

  • The fixed document corpus, including file name or document identifier.
  • Field dictionary with field_key, type, required flag, normalization rule, and allowed values.
  • Optional mapping table for document type, page numbering, table headers, and source-system identifiers.

Result

Write an extraction sheet with document_id, field_key columns, raw value, normalized value, unit, page or section citation, confidence note, and missing or ambiguous status. Add a compact exceptions document for unmapped labels, duplicate candidates, unreadable areas, and schema mismatches.

Quality criteria

  • Every source document has a row or a documented exclusion reason.
  • Required fields are populated only when the source contains evidence; otherwise they say missing.
  • Dates, numbers, currencies, and identifiers follow the field dictionary without losing raw text.
  • Repeated fields retain all occurrences with an occurrence index or line-item row.
  • Every populated value cites page, heading, table, row, or another reproducible location.

Instructions

Apply the field dictionary before interpreting labels. Prefer a value printed beside the matching label over a footer, filename, or inferred value. For conflicting occurrences, retain each candidate and mark the field ambiguous. Normalize decimal separators, date formats, and whitespace only under a declared rule. Never convert a blank, dash, or “N/A” into zero unless the mapping explicitly says so. Record a confidence note based on evidence quality, not on a guessed probability, and keep low-confidence values visible for review.

Adapt before use

  • Add the exact field dictionary, data types, required flags, and allowed values.
  • Map source identifiers, page conventions, and repeating table fields.
  • Define normalization rules for dates, decimal separators, currencies, and identifiers.
  • Set the vocabulary for missing, ambiguous, unreadable, and excluded values.

Related skills

  • Document classification register

    Classifies each document against a supplied taxonomy and records document type, business purpose, sensitivity, retention cue, owner, and rationale. Use for records inventories, contract libraries, information-governance reviews, and knowledge-base cleanup.

  • Document comparison matrix

    Compares two or more controlled documents clause by clause, recording additions, deletions, changed wording, changed values, and unresolved mappings with page or section citations. Use for policy revisions, contract redlines, procedure updates, and version-control reviews.

  • Document translation and localization

    Produces a faithful business translation in one or more target languages while preserving structure, terminology, figures, and formatting decisions. Use for contracts, policies, proposals, manuals, reports, and internal documents; keywords include translation, localization, multilingual, glossary, and terminology consistency.

  • Patientenquittung für eine Versicherungs-Risikovoranfrage auswerten

    Wertet eine Patientenquittung der gesetzlichen Krankenkasse aus und befüllt damit den Gesundheitsfragebogen eines Versicherungsmaklers; offene Punkte werden als Rückfrage an den Kunden markiert.